Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
google.com/goto: Google's anti-scraping update (autom.dev)
624 points by 1e1a 20 hours ago | hide | past | favorite | 485 comments
 help



As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification

I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". It is surprising to see how much they have stripped from our view - long tail results, actual results for product reviews and not ad spam, no preference for 20 page recipe sites.

There are still illegal streaming sports and movie sites everywhere (who knew) and all other seedy corners of the internet that have been neatly erased by Google. It makes me nostalgic for that brief window of time when the web was truly uncontrolled, when page rank had meaning and you didn't know if your search would return 0 results or 4,000 pages, which you could actually browse.


The real "old-school Google, but modern" is Kagi, with the caveat of being paid. (Worth the $5 for me.)

But LLMs can be commanded. This is may bookmark alias for invoking the spirit of old Gog=ogle from within the new, AI-based Google:

https://google.com/search?q=You%20are%20Google%20Search%20fr...

(Cleaned / decoded: 'You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each with a short text from the page, featuring the search terms. Avoid any pages that do not contain the search terms. the search request: %s')


> Worth the $5 for me

Steep at 300 searches per month or about 10 per day. If you search often, it is too expensive. If you rarely search, not worth 5 bucks.

They are plainly trying to push people to their $10 unlimited plan. I would have appreciated if they allowed, say, 600 searches per month for $5 or so.


I find that I rarely do one search. It’s usually a search, followed by looking at the first few links and probably opening them in tabs, then refining the search terms to get better results. It’s an iterative process. So the 10 per day on average would be exceeded quite quickly.

I too feel the yearly price of $108 for one person, $155 for two people and $216 for a family is expensive. Ok, maybe the family plan is cheaper if you can find five others to share it with.


I have the habbit to serach for words to verify their spelling. Very costly habbit when i got kagi.

When i do actually use it for search; it's amazing... assuming I've not already burnt my quota.


I'm not sure if "habbit" was a deliberate misspelling here, but regardless I think it makes the comment better! Lol.

$10 is nothing if you're searching that often to no longer be the product.

https://proton.me/blog/what-is-your-data-worth-to-google


Agree the $10 plan is so worth it as to be a complete non-question as to whether to renew whenever it comes up.

Also, super cool link, thanks for the share.


One of the wildest things to me: Kagi actually warns you before renewals so you can cancel first if you want.

That's confidence in the product right there. I'd argue any subscription which won't do this is designed to exploit people who don't need or use a subscription.


Pretty much sums up why Google’s business model changed. Everyone wants good search, few are willing to be even $10/mo.

You know how many things are asking for "~$10/mo". People don't have the purchasing power they used to, and everything has a subscription model. Firefox with DDG is good enough.

As they said, you wouldn't pay $10/month for search. This is why free pay-with-your-data services won. Nobody wants to pay for what they use, so companies extract value in other ways and you can't complain about that if you're not willing to pay what it costs.

> Nobody wants to pay for what they use, so companies extract value in other ways and you can't complain about that if you're not willing to pay what it costs.

I'm likely old school here, but I never liked "paying for what I use" in the online world because it required revealing some personally identifying information (i.e. to make payment). My attitude is almost certainly irrelevant these days, with the scope of online tracking, sharing of data, and the near necessity of disclosing one's identity to sites other than one's ISP, but it was an attitude that was established in my mind nearly 30 years ago.

(I also don't trust the pay for privacy model. First of all, you can never assume that paying for something implies your data won't be sold. That was even true before the public Internet. Then, even with statements about privacy, you cannot assume that will be true in the future. If the company is sold, management changes, or the company decides the need/want additional revenues, that will change. And that assumes they don't use weasel words about what privacy means, or the ultimate "we reserve the right to changes this agreement at any time".)


Kagi support zero knowledge search through a Firefox extension and Mullvad supports zero knowledge VPN by putting cash in an envelope with your account ID written on the outside.

I mean, there’s more privacy protection on payment data than there is for email address and browsing habits. Plus most companies just push out payment handling to stripe. So most websites have little idea who actually transacted with them to begin with (email dependent).

The US government knows everything though.

Google kept increasing its profits since then, and it was one of the most successful businesses out there way before they changed the business model to become enshittified. The argument that their business model changed because people wouldn't pay for its services doesn't really stand.

If they charged a newspaper price for it per se (like $1 a month), Google would be rolling in cash with as many users they have.

Newspapers are not $1/month.

And Google's ad revenue is likely more than that. Or, if it's not for every user, the ones who would opt out of ads are the ones who can pay. You're cannibalizing your highest value users.


You are worth much more than a dollar. Each search is worth between 10 and 90 cents. If Google shared that with the user people would be happier

I do agree it's kinda expensive. I'm lucky enough that my employer pays the $25 plan for met because it includes access to many LLMs.

I pay the $10 and use assistant multiple times a day and never needed the fatter models since for coding I have copilot and Gemini

True. I got this as my first LLM subscription (was a long time skeptic, less so now but still a bit), but now I have cursor too so maybe I can even downgrade. Thanks for the reminder. :)

> because it includes access to many LLMs.

That is interesting, which LLMs? Asking as someone paying monthly Claude subscription.


A lot of them. Claude is a choice. Here is the list : https://help.kagi.com/kagi/ai/assistant.html

You can even choose a model for each turn.


My time, attention, and “brain” are valuable. Having aligned incentives with my search provider and control over the results it serves me through weighting is worth substantially more than I pay for unlimited.

Kagi is the stickiest digital subscription I have.


It's relatively rare as of late that I need to specifically find something on the internet, and DDG does not suffice, so I rarely exhaust the 300 Kagi searches.

Most of the time I need to find out something, and in that case, Claude / Gemini / Grok / Perplexity / you name it work better. Most of my queries are technical by nature, not interesting for data mining on me personally.


The 300 searches a month work for me because I only use Kagi when I think it will give better results than others. And most searches are trivial so any free engine will do (like you have the website name and want the url).

The $10 plan is worth it.

FWIW Kagi has worked hard on their pricing over the last few years and has been trying different models. They're not "big search" the price is what it is so it can exist as a business. 100% fine to not be a customer obviously, but then you forfeit your license to complain about Google spying on you and ruining the web. Like they're trying to solve the problem. IMHO speaking just for myself I feel a moral duty to support them.

You can not find the value of 10 searches a day for 5 dollars and complain about Google. Yandex,DDG and Bing are major search engine choices.

Kagi has limitations like not showing sites with ads which many people with ad blockers don't mind.


I agree on all your points. I'm just saying, that the $5 tier is not a honest to good option. The $10 one is but they can make the $5 one genuine too by increasing the quota. I'm sure they have done their research and maybe increasing it will cannibalize the $10 tier.

Good workaround! It's just sad that to achieve the previous behavior we now need to burn significantly more compute, and in turn energy, and with far worse performance and an inverted UX.

The most frustrating part is that they have all of the data, and oodles of compute, available to surface the same very functional experience they originally offered, but they would prefer the image of being visionaries rather than the reality of being useful.


I think the "original experience", whatever it may have been (and I strongly suspect there's more nostalgia here than actual worse results) would be very quickly overrun by modern SEO methods / tactics.

Kagi can be "better" than Google (for some very narrow definitions of better) because they basically don't have to think about SEO at all. They're small enough that nobody does SEO against Kagi. If their algorithm is sufficiently different that typical Google tactics don't work on it, even if it's much more naive because of that, Kagi wins.


Thanks, I needed to add a ?aep=11&atvm=2&udm=50 for AI mode. Wow!!!!!

>You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each with a short text from the page, featuring the search terms. Avoid any pages that do not contain the search terms. Do not simulate fetching results. The URLs must be complete and not truncated. No markdown. The current date is 2026. The search request:


Ah haha that is kind of funny it finds exactly what I was looking for, just much slower than 15 years ago.

I'm using Kagi every day, and unfortunately, it is not true anymore. I think what happens is that the blanket banning of IP ranges because of LLMs, and flat out incompetence and laziness, affects them heavily. I get results, or just a few unrelated ones many times every day, while Google happily returns what I searched for. I think Kagi's index is shrinking, or they really made their algorithm worse.

I slowly shifted to use LLMs because of this in the past half years, making the situation probably worse for Kagi. They were the last usable search engines for years now. I still pay for it, but it's increasingly not that useful anymore.


Same here. Been paying for Kagi for probably a year or so, but I find that it hasn’t really been able to keep up with at least my personal expectations. It feels quite slow (which I don’t remember it always being), and the search result quality is not always the best. On top of that, the 10$/mo is a little steep for me, given that their assistant and maps are quite sub-par in my opinion. I would have preferred them to focus more on the search experience, and less on the other products. Or give me a cheaper plan for only unlimited search without the other products.

> It feels quite slow

It definitely is. The extension-based redirect in Safari is also unreliable. Furthermore, many of my searches (and others', apparently) unexplainably defaulted to Groningen in the Netherlands, even though I'm nowhere near the place.

I ended up not renewing my annual subscription. Since then, I vibecoded a small userscript that replicates domain ranking for DuckDuckGo – plus DDG already lets you turn off ads natively.


I like Kagi, but is it me or have their results been less than great for the past 1-3mo? Maybe the places they're pulling their results from have had some changes?

I haven't noticed.

> You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each with a short text from the page, featuring the search terms. Avoid any pages that do not contain the search terms. the search request: %s

I paid for Kagi a couple months, and was rather unimpressed with the quality of results.

That prompt mixes in unrelated results about the workings of search engines. You could just use `udm=web`:

https://www.google.com/search?q=%s&udm=web


Kagi seems to heavily rely on Yandex for results. I don't use it, so can't compare, but are results much better than using Yandex directly?

I'd say that the ranking, and thus the relevance, is noticeably higher. If I know what I want to find, Kagi is good at surfacing exactly that, not something vaguely related.

I see, thanks.

Kagi is good but it’s about 25x too expensive (no really).

Oh that shorthand is fun. Curious it works at all.

But yeah I'd switch to Kagi if I had any more forcing function. Maybe without LLM's but google continously getting worse, that'd been the actual thing for me


Kagi works by scraping Google and other engines, so this should still worry you.

I still support the use of Kagi though, as a market signal to Google that we'd even pay them for their product if it wasn't dogshit.


For Google, the business side of you paying them doesn't work out as obvious as it might seem.

Advertisers are sold the idea that 1000 clicks/impression is X dollars. The average impression might be priced at 1000/X$.

But you are worth much more than that for two reasons.

- you are a person willing to pay 5$

- Google can sell a “collectivized” product. The same way that eg collective farmer crop insurance creates actual value.

and the third reason (as consequence of the second) they dont want people to think about: it gives them more opportunity to fudge the value by making it all very complicated.


The advert paradox, anyone willing to spend money to avoid ads is inherently worth more to advertisers, since they're willing to spend money.

The more money you're willing to pay to avoid ads the more you're worth to advertisers.

That's why tools like uBlock and YouTube morphe are the only answer, negotiating with terrorists never works.


I'm willing to pay time or money to avoid ads, not to be spied and harassed by companies making me offers to buy something that almost certainly I don't want to buy. Those companies should pay me a percent of what I make them save by not buying ads to serve me.

And eventually when I want to buy something I research it and buy it and nobody had to pay anything to Google.


> That's why tools like uBlock and YouTube morphe are the only answer, negotiating with terrorists never works.

This is silly. The economy doesn't work if no one is willing to pay for goods. The content you're consuming won't be made if theres no market value to create it in the first place. Pirating is not a solution to our data privacy problem.


The capitalist economy is based on the explicit idea that everyone will do everything they can to extract value from everyone else, and it'll come to some kind of equilibrium. If you do not do this but you allow everyone else to do this, you receive worse quality products for ridiculous prices and do nothing about it.

In other words, you owe me $1000 as a condition of reading this comment. Your options are to unread the comment, pay me $1000, or pay me nothing to incentivize me to charge a price you'll actually pay (likely $0.00 for comments but not for everything)


I don't understand your logic. Why would someone that never sees my ads be valuable to me? Why would I care? Everyone spends money on something, the fact you spend it on adblockers doesn't make you automatically more valuable to advertisers. I have never heard of this "advert paradox" concept you mention and I have worked in industry for a while now. All this is just conjecture on your part and doesn't really correspond to how it works.

If you have $20 to spend avoiding ads, you have $30 to spend on cool products so the advertiser is willing to spend up to $24 to make you not block ads about a $30 product you'll buy (that costs $5 to make)

> If you have $20 to spend avoiding ads, you have $30 to spend on cool products

Yes, but if you are so against ads that you spend $20 to not see ads (I am definitely in this group) and then see an ad for a "cool product" anyway, then that product is no longer cool. Any time something like this happens, I legitimately develop a grudge against that company and their products.

I might have bought that product before seeing the ad, but there is no chance of that after seeing the ad.


That does not make sense at all. Why would you care about people who want to avoid ads? The real data shows very little people actually do that. How much you spend on adblocker is irrelevant to your profile for a business. They will get you some other way if you ad block on digital.

Sad but true. I paid for Prime to be ad free, easily worth 3$/month but sure enough now part of their catalog has become "only available with ads" or ads in their FireTV is before you hit the prime app (and then again when you do). So now their whole product is crap and I'm looking for a new streaming stick solution [anyone?] Sounds like it's time for a newcomer/disruptor like Angel Studios to release one.

Natural next step of the ad lobby is obviously to pay the lobby to make it illegal to bypass the money making machine if they haven't done so yet. I don't doubt that capitalism still works, it just takes time to "boil the frogs" dead [enough] to stop them from buying such crappy products.


If they don't want to sell me a deshittified product, I'll continue paying someone else to extract the deshittified product from them. Their choice.

Kagi don't scrape, they pay other engines for API access: https://help.kagi.com/kagi/search-details/search-sources.htm...

>Our search results also include anonymized API calls to all major search result providers worldwide


The only “other engine” that matters is really a thin paid SERP façade in front of a scraped Google result. In other words, Kagi pays for access to an API whose implementation is basically live scraping Google and throwing away the ads.

There are a few other indexes that Kagi’s aggregator mixes in (such as Bing Search) but those haven’t been contributing much in terms of meaningful results. The sheer size of Google’s index still dwarfs all the others, including Bing Search.

So yes, Kagi doesn’t scrape but they pay SERP providers who do.


Marginalia likely returns better results than Google, when it returns useful results. DDG/Bing likely returns average quality results which can be better than Google's slop results. It's not true that Google is the only useful one.

They pay scraping sites for results, including SERP API. When I pointed this out, last time Kagi was discussed on Hacker News, an employee of Kagi said that they're trying to build their own internal index, but he didn't provide details.

Is Marginalia the only real competing search engine?

AFAIK Brave Search is also using its own index.

I’m pretty sure Kagi’s entire pitch is that they’re an alternate and original dataset. Theirs is the only search engine whose data isn’t based upon google or bing’s databases.

To anyone who didn’t find Kagi useful in the past, I strongly recommend you try it again, as they have made huge strides in usability and quality of surfaced data in just the past year…. Huge strides since even a few months ago too, with more adaptive curation, data labeling, and UI/UX


> I’m pretty sure Kagi’s entire pitch is that they’re an alternate and original dataset. Theirs is the only search engine whose data isn’t based upon google or bing’s databases.

It's not. In addition to their own indexes the results include calls to "all major search result providers worldwide"[1].

There are independent search engines that only use their own index, like Mojeek[1]. Kagi is explicitly something like a search engine aggregator.

[1]: https://help.kagi.com/kagi/search-details/search-sources.htm... [2]: https://www.mojeek.com/


Kagi is mostly based on Bing and Google results, officially by API and not scraped, and ranked by there own mechanism I think.

Kagi was using Yandex, it comes up on every Kagi post here. Seems that they've removed that info from their site now though...

Officially Kagi pays SerpAPI to scrape Google for them, since Google doesn't have an API they'll let Kagi use.

Interesting. So this means the new google.com/goto will exclude Kagi from the results?

Not necessarily, as evidenced by TFA, whose scraper still seems to work even after it visits each individual SERP entry to harvest the URLs.

But Kagi’s SERP API provider is going to have to adapt, I guess.


...you.. do know classic Google Search is still an option, right? Just add `&udm=14` to the search URL.

Also Google AI mode is accessible on google.com/ai.


+1 for Yandex. I was not expecting it at all but agreed you get most of organic search results which is what search used to be based on relevancy.

> For actual web search I actually like using Yandex

You might want to reconsider this.

https://meduza.io/image/attachments/images/007/699/276/large...


Absolutely. Where google is shaping results "to please corpo-political masters", Yandex is obviously influenced by the Kreml

I use Kagi as my daily driver. But there are some kinds of queries where my interests and Yandex's interests happen to align. Mostly anything corporate America wouldn't like, like evil copyright-infringing websites


I sort of agree with you.

Meduza gonna Meduza. Go do the search yourself. I get nothing like what Meduza claims.

Yandex censors the results, it is required by law, so I do not understand what are you arguing against. To be specific, any URLs, which are blacklisted and banned in Russia, must be omitted from search results. Which includes BBC and other Western media and explains the difference in images because in the Google's results the images come from BBC and Voice of America.

Also, if you try to search for "download Chrome" (in Russian) then the first result in Yandex leads to a scammy website: https://ibb.co/HDRJGgZD The real link is the second one, but I remember a year ago or so there was no official link at all. You can also note that official link to Google has a grey text saying: "the owner of the resource violates Russian law" (Yandex is required to show this notification).

Yandex has also been caught "accidentally" removing the site of Ekaterina Duntsova who was planning to nominate for presidential election in 2024, and showed fake/phishing sites instead.

So, Yandex is only good for searching torrents/pirated movies (which you can watch on rutube) and nothing more.


Yeah I did while I was in Kiev at the time. Did not screenshot unfortunately but yeah that's what Yandex did show for many "Bucha" search terms.

You think Google doesn't do censorship? Or any mainstream search engine, for that matter?

It's the usual "I prefer censorship not targeted at/tailor made for me".


Bold to assume Russian propaganda is not aimed at the west.

[flagged]


Naivete and scorn of autism in one short sentence - you're very efficient.

Yandex.com is the international front end to yandex, and not a portal for 'a handful of western autists'.

But on Russia's granular propaganda and disinformation efforts, you ought to read Peter Pomerantsev's excellent "This is Not Propaganda" before making silly sweeping statements.



surprising to see how much they have stripped from our view

Don't forget cached pages.


You may not even appreciate the full extent of it. Long ago now, when Google took off and started dominating because it was “the best results” people didn’t realize that even though they were the best results, they were not actually good results. To be the best, you just have to be notably better than your nearest competitor, you don’t actually have to be good or even decent, just noticeably better. Today Google really has no competitors besides all the derivatives it has maintained to maintain an illusion of competition, the most obvious examples being DickDuckGo and Bing.

It’s somewhat similar with AI today, only the competitive landscape has majorly shifted with the Chinese models, something that wasn’t supposed to happen. All the sudden you no longer have a closed system of American models that can all be contained and managed as good enough to make people believe are “the best”; they actually have to compete in a far more competitive arena with the tension of the best model being the most competent model, and the most competent model, by virtue of AI, will be the least restricted, the least censored, the least guardrailed, the least controlled to present a telescreen and Hollywood type fantasy world where the good guys always win and … wouldn’t you know it … the good guys is always us, the people controlled by a psychopathic, narcissistic cabal that also controls what AI will tell you, the false truth.


It's a shame there aren't any good Chinese search engines. From what I've read the Chinese internet is mostly centralised around a few apps like WeChat.

A nice story, but I don’t think there really was a time where closed source models’ development stalled really.

Which being public wealth, should be publicly available.

Data that were sourced from the public, should be public. We need this legislated, or wealth inequality and fiefdoms will keep growing with the historically known outcomes.


Didn't Google take it down because of copyright? With the rise of paywalls it became a backdoor around them, but Google doesn't actually have a license to republish the copyrighted text they scraped.

It's not a backdoor is the sites intentionally served Google a version without the paywall so the indexing got "better" than what real users got. The sites kind of dug their own hole here, Google wouldn't even sit on the content unless it was offered up to them in that way.

Unfortunately copyright law allows them to give something to one person without giving it to others _and_ to sue to the former person to block them from giving it to the others.

Just tried Yandex again, after a long time, and indeed it was a blast from the past, in a positive sense! So many real search results, which transports the feel of an authentic mirror of the web contents. Default search engine from now on! Thanks for the hint.

Yes, it’s been impossible to trust Google since the “Jigsaw” project, and especially when they started blanking out search pages on trending topics with messages like “these results are changing quickly”. It’s clear that they have moved from being a neutral arbiter to pushing a perspective.

I refuse to use a search engine that sees me as cattle to be steered towards “preferred” information sources. I specifically enjoy reading fringe material and will continue to seek and find it, thank you very much.


> There are still illegal streaming sports and movie sites everywhere (who knew) and all other seedy corners of the internet that have been neatly erased by Google

This is such a strange position to take. Apple doesnt allow all sorts of apps on the appstore, but that is never said as "Apple is erasing illegal streaming". Google is simply not showing the results. They are not taking down, or banning, or doing anything to the websites.


The pain being made is: If you use a search engine as your eyes to see what exists “on the internet”, then, absolutely, whatever Google hides from its results or fails to index is “erased” from “your” experience of the internet.

I see where you are coming from but the argument would be stronger is Google Chrome refused to open illegal streaming sites etc. AFAIK, there's no such restriction.

(I dont agree to this but ...) using your argument of " If you use a X as your eyes to see what exists “on the internet” .. " - we should all be mad at Apple. I use iPhone as the primary device to access apps, and they not just hides but actively ban and cut whole swathes of developers.

If I ask Siri to give me a link to illegal streaming site and if it refuses, is that cause for concern? I'd say no. Infact, I dont expect it to give me that and I get it. Same for Google Search in my humble opinion.


> I see where you are coming from but the argument would be stronger is Google Chrome refused to open illegal streaming sites etc.

Why would it be stronger? Google has a near-monopoly on search.

> If I ask Siri to give me a link to illegal streaming site and if it refuses, is that cause for concern?

Yes, because you cannot replace Siri on a device you own.


I see piracy sites just fine on Google. Almost always the top result. That's why I go there to find the next domain after a previously working one gets shut down.

If you use an iphone then yes, apple is indeed "erasing" illegal streaming in the same sense that google is. We can be nostalgic for the old internet while also recognizing that at least google is well within their rights not to serve up results that fall outside the bounds of the law. (Apple not so much. When you gatekeep the hardware platform I think you're ethically obligated to act as a common carrier. Unfortunately the law doesn't require that.)

> Apple doesnt allow all sorts of apps on the appstore, but that is never said as "Apple is erasing illegal streaming"

Because they are not "erasing" anything, they are refusing to provide a platform that enables the direct (app whose intended purpose is) delivery of illegal content to you.

The difference being one will not facilitate the activity, while the other, is actively suppressing it. And thats what is meant by erased.

Your question in other comment If I ask Siri to give me a link to illegal streaming site and if it refuses, then that would be the better comparison. Then we could say "other seedy corners of the internet that have been neatly erased by [Siri]"


If the companies weren't trying to destroy trust and take everything away from everyone these things would never have much traction but now they deserve to have more than ever.

>I do enjoy using their free AI

It's really not free. You're paying with your data.


People have this fixation on privacy, when there's so much more going on, potentially way mor important.

Google's strong position in AI gives them bragging right to attract more companies and get them to actually pay, subsidizing your use. It also lowers competitors position as you're not touching them while we're on Gemini. It also fortifies their position in the future ad market.

Being second or third in AI usage is worth a lot, one's private data matters very little in comparison.


It's two totally different matters.

I won't be first, second or third in the AI race. I'm not even participating in that race. I don't care who wins. However my private data are mine and I do care about them.

It's Google and those other companies that may not care about my privacy, because the AI race is more important to them.

If AI is going to cost too much, I will use a cheaper one (Ferrari or FIAT?) or if every AI will cost too much, no AI all and there will be billions of people like me.


> You're paying with your data.

Parent was pointing at the value exchange, and I argued that users' privacy is not the currency Google deals with.

Your privacy can be valuable to you, it just has no weight on the exchange.


> Being second or third in AI usage is worth a lot, one's private data matters very little in comparison.

what does this even mean


Worse - We all are paying with our planet.

what is that not true of? HN and your post on it is paying with our planet

Everything uses energy. AI is uniquely bad due to its scale. Comparing AI inference and training with posting a one-line comment on a website is like comparing wildfires with candles because they both produce heat. It's true but useless as an argument.

Complaining about AI is just a meme. There are tons of things that use way more energy for arguably way less utility to humans.

AI was not even a thing a few years ago so almost anything that use more energy than AI is probably far more useful

Plus soon we'll be paying based on secret parasocial influencing agendas.

It's now possible to put a complex spin/bias on whatever the system shows (or whatever it chooses to bury) in a really easy and scalable way.

Ex: Dairy Association pays, and suddenly results about bones are just a bit more likely to show something about the importance of calcium that many people get from milk. Fraternal Order of Police gets involved, and now "shot by police" everywhere gradually morphs into "was in an officer-involved shooting".

... And of course "Google making it harder to see URLs" becomes "Google taking bold steps against evil scrapers."


I have an android phone, they already have my data.

So have I, but they don't have any of my data.

Me too. I don't connect my phone to my Google account. I understand that the Play store is enticing, but there are other options. Some of which are actually better.

I made https://froogle.fyi out of a desire to get that old school search back. Source: https://github.com/scosman/froogle

It hasn’t replaced Kagi as my daily driver but the UX is fun.


“Froogle” was an actual Google offering (years ago) that now redirects to Google Shopping.

Sometimes I use my own index of domains and channels

https://github.com/rumca-js/Internet-Places-Database

I hate the very idea that destination location is opaque, and user can't verify if you are funneled toward malvertizing.

I hate even base64 encoded links in the results.


This is really neat, thank you for sharing.

Yandex also easily finds copyrighted stuff and other various types of wrongthink banned by Google.

This is interesting. Do you consider Yandex absolutely safe to use (wrt their possible connections with the Russian government, and current war in Ukraine)?

Asking this as a geeky end-user from Estonia -- I wonder if I would be considered "suspicios" by local internet service providers, our govt/police etc if I did a lot of searching via Yandex instead of e.g. Google these days.

(For us in Estonia, I suppose "everything with possible Russian ties" seems suspicious these days, unfortunately with a reason. As a side effect, this also creates a hesitation to look into interesting projects like ReactOS, the old-dos.ru repository, etc. Russia, obviously, inhabits lots of extremely talented coders who still have the skills and mindset to push older hardware to its limits, actually still write stuff in assembly for that, use the reviving DOS distros, etc -- lots of very interesting, creative coding done there, I guess. But, I hesitate visiting their webpages due to Putin's aggression in Ukraine, them being pro-Putin, and possible spying related to all this. It's a sad state of things actually. Sort of like a loss of computer-cultural ties.)


I’d expect everything that you type into Yandex that can be of use to the Russia will be used by Russia - they will nit care about hurting you.

Just like Russia doesn’t care about online criminals being located within their borders - as long as they target outside.

While this may be true to some extent of all the countries, Russia is the one to be clearly against Estonia, Baltics and eastern europeans.


How is that different for any web site on Earth?

As a general rule, if another country has an open war with your neighbor and openly says that you may be next, don’t put in private data into the sites controlled by that country.

But that's true for google too

Ok, maybe not the war *with a neighbor* (yet), but US has an open war with Iran (and has bombed a few other countries recently also) and it also threathens to annex the neigbor to the north plus another island country.

Sure, americans might consider themselves to be "the good guys", but the rest of the world doesn't.


From where in the rest of the world are you? Did even you talk to anyone living in countries neighboring Russia?

Sure not everyone benefitted from US but a ton of people/countries did and still do.

US never felt a need to build walls to prevent their citizens/allies from leaving. Russia and the others very much so.

With US the track record may be mixed, but a ton of countries from Europe and Asia benefitted tremendously. Russia otoh doesn’t have a concept of win-win, they tend to exploit even their closest allies, which anyone living in Baltics/East-Central Europe can tell you.


What bombed country benefits from US? Afghanistan? Iraq? Iran? Yemen? Venezuela? Syria? Yemen? How do they benefit? Destroyed infrastructure is somehow helping them? Americans stealing their resources if they decide to occupy the country? How? Win-win in iran how exactly? US just proved to the rest of the world that you need nukes to be safe from americans, and even then they'll sanction you to hell if possible, just because you're not bending over for them.

I'm from the border of central europe and balkans, I don't benefit at all, it just costs me money to pay for a van full of soldiers every time americans decide to occupy some country half the planet away.


Private mode. Tor browser

anyone from Ecosia gang?

How is this comment relevant to the post (and be so upvoted?)

Google is making it harder for people to game the index, working in your favor.


Mmmh this change doesn't seem about gaming the index but rather about scraping the index. Which affects in a negative way only Google. And actually would affect users in a positive way because it lets other companies create their indexes more easily (at the expense of a poor mega-corp, indeed)

Are they? This reads more like a way to stop people reusing their results in other services rather than reducing spam results, but I may be mistaken?

Have you tried Kagi?

> I do enjoy using their free AI.

So how is that AI slop garbage useful? It often lies, aka "hallucinates". And I operate within this Google-controlled auto-spammed text of lies when I use its horrible AI. The last thing I want to do is make Google even more powerful; at the least now with Google search being so crap, alternatives would be well desired, but oddly enough I also get crap results using e. g. DDG, so I think in part the whole www stack must have become more crap too, quality-wise.

I actually don't see it as there are extensions that eliminate this slop spam, but I never found it useful for anything. Other than waste my time, back when I still saw it. Thankfully I no longer see that AI slop spam due to these extensions.


> Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters".

As long as you don't search anything related to Russia itself, or to Russian interests elsewhere (like their invasion of Ukraine). Then it's very heavily biased, priority is given to state ran propaganda mills, independent media is hidden from results, etc.

And before anyone starts thinking about whatabouting, Yandex are based in a country where journalists are openly and publicly assasinated to intimadate. It would be delusional to expect any sort of press and related (like search engine or aggregation) freedom, or to try to compare this to anything in any other developed country.


Reminds me of this old joke.

A Soviet citizen and an American are sitting next to each other on a commercial flight.

The American turns to the Russian and says, "I have to hand it to you — your state propaganda is very impressive. You really know how to shape what people think."

The Soviet smiles, nods, and replies, "Thank you, but it's really nothing compared to American propaganda."

The American looks shocked and says, "What are you talking about? We don't have any propaganda in America."

The Soviet smiles again and says, "Exactly"


Nobody is denying the existence of American propaganda, but it's an absurd false equivalency to pretend it's comparable to Russia - the country where there are multiple Wikipedia articles with lists of murdered journalists and there is a day of remembrance for murdered journalists. The country which sends assassins to kill dissidents abroad. Especially when we have the Mitrokhin archives since the 1990s in which a KGB archivist painstakingly tried to warn us about the firehose of shit method. Do you know the conspiracy theories about AIDS were started by the KGB? This is the kind of shit we're talking about, not Americans jacking themselves off on their anthem thanking troops for their service.

Actually sorry, it's beyond absurd.


> Actually sorry, it's beyond absurd.

So is Fox News. That's the OVERT propaganda.


So, I decided to check, and the first links that show up when I search for "war ukraine" in yandex.com are Google news are from rbc.ua and bbc.com. When I search for it in yandex.com, the first link is CNN, the second google news, the third is indeed TASS, but the fourth is novayagazeta.

I then picked "bucha" as the query most likely the be affected by propaganda and censorship. In yandex.ru, there are propaganda outfits in the result (something called ruwiki.ru comes third), most links are of fairly direct accounts of the massacre by independent media. The AI summary says the town was "occupied" by the Russian army and "liberated" by the Ukrainian one and "После отступления российских войск в городе были обнаружены многочисленные свидетельства массовых убийств мирных жителей." On yandex.com on the other hand, the first page of results does look fairly propagandistic, including "globalresearch.ca" and "donbass-insider.com", which seem like propaganda outfits. But it also does include accounts from novayagazeta, al Jazeera a video of killings from Radio Free Europe.

While they do put either own propaganda, neither Western, nor Ukraininan nor independent media is not hidden from the results and it's not visibly de-prioritized.

Of course, I wouldn't discount the possibility that the results would look quite different from Russian territory. And, if anything, half the reason why Russian propaganda is so effective is that its self-aware and capable of subtlety when needed. "Fuck you if you can't handle the truth, this version of Biden is the best version ever." would never happen there. Unlike Scarborough, Soloviev knows exactly what he is.

Open and constant suppression of information is not the regime's usual strategy in information space (though, they'd ratchet up the level of control when they deem it necessary), which is why I knew the claims here wouldn't stand up to scrutiny.


Yeah, the Bucha search was a perfect nail in the coffin of previously very respected company: https://meduza.io/image/attachments/images/007/699/276/large...

Have your tests been ran on vpn with vpn cloaks?

I explicitly say, that I won't be surprised if the results change from Russian territory. I already put way too much effort to check the validity of an internet comment, but if anyone is interested, they can check.

I wish yandex worked well, but the enshittification wave reached it too. It used to have such a great reverse image search, now it's useless

Any proof of that? I’ve always found Yandex to be substandard but never for reasons of censorship.

there was a source code leak and you can see some examples there which make the president's image more alluring. e.g. if you search for "Putin" they add an invisible "-crab" to your query

that's the most innocent example from the top of my head. There probably are many more serious censorship issues


Yep, plenty, and tbf it's entirely logical. The people working at Yandex do not want find themselves dying of a nerve agent or polonium, which is a real and acute danger for people who displease the ruling regime.

https://hal.science/hal-03217497

https://euvsdisinfo.eu/yandex-from-tech-innovation-to-inform...

https://pmc.ncbi.nlm.nih.gov/articles/PMC10130930/

https://misinforeview.hks.harvard.edu/article/a-story-of-non...


If it's true just link to a first party example on the yandex site

[flagged]


What does israel have to do with this?

According to OP

>It would be delusional to expect any sort of press and related (like search engine or aggregation) freedom, or to try to compare this to anything in any other developed country.

Israel is responsible for killing more journalists than any other country. Google has profited by providing services to Israel in the ongoing genocide.

https://cpj.org/special-reports/record-129-press-members-kil...

https://afsc.org/gaza-genocide-companies#Google:~:text=Googl...


I was waiting to see how long it took for someone to "make this about Palestine" - their strategy is to inundate all discussion with the topic.

Over half the journalists in the Gaza Strip are affiliated with Hamas. Many of them infiltrated Israel in 2023, and some of them held Israeli hostages in their homes, including children hostages.

https://www.terrorism-info.org.il/en/abou


So Israel killing "bad journalists" is fine, but Russian are killing only the good ones?

Nice world to live in.

You should know that any country can declare a journalist a foreign spy, a terrorist, but does it give them right to kill?

Without due process, without trial? Because someone says "they are Hamas".

It's Godwin time.


Hamas claims that the journalists are Hamas, not Israel. And I'm not declaring any journalist "good" or "bad". I'm stating that a Hamas member is a valid target. Him moonlighting as a journalist does not change that.

Every claim need to be proven and everyone deserves a process. In a democracy of course, it may be different in a fascist state committing a holocaust.

That's not how things work when you are participating in hostilities in an active war zone. What a ridiculous statement to make.

That site seems completely unbiased and not at all fed by the propaganda coming out of Israel. Thanks for finding such a reliable and trustworthy source.

It only counts when the country is our declared enemy. When our ally does a holocaust, it isn't a holocaust, it's just anti-terrorism operations.

Your Holocaust inversion is not only morally wrong, it is also factually incorrect.

Care to explain?

Parent comment is implying that an ally (Israel) is performing a holocaust. In fact, there have been a good dozen genocides in the past century which I would agree fit the term, but not the conflict that Israel is involved with.

The anti-Israel and anti-US side typically take an argument or incident from Israeli or US history or culture and invert it. A good example is the common Hebrew expression "No other land" - we have culture and songs about that, as well as jokes and it is a common expression in everyday speech. After October 7th, as part of the propaganda campaigns against Israel, that phrase was used as the title of an anti-Israel movie highlighting that another population has "no other land". Inversion of our culture.

Holocaust inversion - accusing Israel of perpetrating a holocaust - is this particular example.


As far as I'm aware, when a country sets out to exterminate* a race, that's a holocaust. Am I wrong?

* Often the torture is the point, and they don't completely eradicate the race because then they wouldn't be able to keep torturing them


Israel never set out to exterminate, Israel set out to get her hostages back - including babies under a year old.

You know who did set out to exterminate? Hamas, who in their very own charter state their intention to eradicate the Jews: https://avalon.law.yale.edu/20th_century/hamas.asp

I suggest reading this article which explains the propaganda campaigns that followed the October 7th attacks on Israel: https://www.adl.org/resources/article/hamas-its-own-words

Only one side is calling for genocide, and it is the Arab side - propaganda and disinformation campaigns aside.


Please make the same analysis for Germany in regard to the Gleiwitz radio station

To other readers: Parent is suggesting that the 6000 Gazans who breached the Israeli border on October 7th 2023 and killed 1200 people, mutilated bodies, beheaded humans and dogs alike, raped women, and took over 250 hostages including infants... Parent is suggesting that these people were undercover Israeli agents performing a false flag operation.

A genocide with an increasing population. And a famine where obesity is rising.

Jew haters on the internet?

Never heard about it.


> And before anyone starts thinking about whatabouting, Yandex are based in a country where journalists are openly and publicly assasinated to intimadate

Thankfully Gary Webb is here with us today to laugh at this


[flagged]


Have you tried Yandex and evaluated the quality of its results, or do you just feel obligated to continuously put it down because it's russian?

I turn it down straight away for being Russian and I'm suspicious of anyone that mentions it positively.

Russia is in a cold war with democracies in general and a hot war with Ukraine. It has an autocratic dictator as leader and state control of media.

It has a well documented propaganda campaign with the goal of destabilising other countries through agitation. We even know the address of the building in Saint Petersburg it's run from.

Why would anyone try to assess Yandex's results? They could be like a slot machine serving good results 78% of the time and slipping propaganda in when least expected.

Why bother with that when there are good, free alternatives?


Couldn't we say the exact same thing about any US service in this day and age?

Trade war with half the planet, hot war with Iran. Autocratic dictator as leader and complicit media. Well documented propaganda campaign and history of destabilising countries around the world. Known to compel tech companies to spy on foreign citizens...


[flagged]


It's not whataboutism, it's calling out a "proving too much" fallacy https://www.lesswrong.com/posts/G5eMM3Wp3hbCuKKPE/proving-to...

Maybe so they can bring actual data to the discussion instead of bias.

I have, and it has never been good for my purposes. In English, I just can't find things I need, in Russian, I find things I wish to never see again, tbh.

What is your experience with it?

Edit: The best search for my purposes historically has been on Pinterest. Unfortunately, not anymore. I normally search for visual artifacts (photos, art, schematics, etc), out of all search engines Google is still the only one that is somewhat functional for me in this regard.

I tried DDG, Kagi, Yandex. None of them worked for me.


Have you tried its image search? It's pretty good compared to its competitors. Google's image search fell off as they severely nerfed it.

Google is terrible these days, I agree.


How about topics the Russian government doesn't care about?

How about topics that USA censors?

Of course you should not search Ukraine info on yandex. But you are only showing the worst case scenario. What's the average and best case?


As long as there are non-russian alternatives out there I'm never going to even try.

> campaigns for a Russian product on HN are annoying, and appaling if you have any morals

How is Yandex more morally-bankrupt than Google in your analysis?



The quintessential HN comment questioning the obvious.

Yandex does Putin's bidding much much further than any western search engine ever meddled with their results for political purposes. Yandex employees' lives would be actively at risk if they didn't.

Sometimes I really wish a hard reality check would meet the people making such comments from the comfort of their programming chair or from thousands of kilometers away.


> Yandex does Putin's bidding much much further than any western search engine ever meddled with their results for political purposes. Yandex employees' lives would be actively at risk if they didn't

Does media manipulation score higher/lower on your moral index than the widespread domestic spying the US has forced tech companies to be complicit in?

Not that media manipulation by the US government is exactly an unheard of event (at least as far back as the whole WMDs in Iraq scandal).

> Sometimes I really wish a hard reality check would meet the people making such comments from the comfort of their programming chair or from thousands of kilometers away

Look, I'm American myself, and have as keen a dislike of Russia's foreign and domestic policies as the next guy. But it is at best wildly hypocritical to suggest that US tech firms are not subject to government interference, in a manner that may be detrimental to their customers.


> But it is at best wildly hypocritical to suggest that US tech firms are not subject to government interference

It is very different degree of subjection where we talk about country that have laws, politics, сourts, journalism, all that stuff, and country, where if you dare even to complain about some wild highly illegal stuff that the kgb demands corporation that you working for to do, you end up getting 30 years in prison for treason, falling out of the window, getting poisoned by some chemical weapon, and even then you will be considering yourself a lucky man because nothing very, very bad had happened to your family.


Guidelines: Please don't use Hacker News for political or ideological battle. Please don't post insinuations about astroturfing.

[flagged]


> There are "politics and ideology" and there is right and wrong.

That is exactly how six-year-olds think.


Thankfully I'm not a 6 year old, and you're not my dad.

If anything, yours is the kind of retort a 12 year old would find clever.


I'd only use it via tor or private mode behind a CGNAT or VPN

There's no doubt about that but it definitely meets some people's needs.

IMO only useful for piracy.

I prefer Google itself but this goto update is terrible.


Do you use any Israeli products? What do you say about them?

I mean, if I were to accuse any search engine of an astroturfing campaign on HN, I'd certainly not pick Yandex over Kagi.

I don't understand why people are still using Google search generally. It's so bad nowadays.

I did an interview with Google around 20 years ago, where they posed a challenge involving tracking which specific search results people click. It's obvious in hindsight the solution required rewriting all the urls to redirect through their servers. Note this was in the days before they already did so as a matter of course.

I failed to gain traction on the problem, because to me the very idea of doing such a thing was too reprehensible to seriously consider. It broke an unwritten contract between the company and the user's expectation of how websites worked. You expect to be able to do things like right-click a link and copy the authentic URL, or hover to see where it wants to take you. The notion of obfuscating the link beyond easy recognition and polluting it with tracking markers felt misleading and, well, evil. A move that would mainly only benefit Google, and not it's users. I (quite mistakenly) presumed this opinion would be obvious and self-evident to anyone who spent enough time around the early web to understand its norms.

I explored other ways of achieving the goal, but it clearly wasn't the answer the interviewer sought.

I'm more seasoned now, and experienced enough to say with confidence the approach was wrong. This may seem like a small thing, but a series of misteps and chronic failure to adequately advocate for users is what has led us to the toxic waste dump that so much of the Internet has become today.

I'm really glad to have fresh alternatives (like Kagi), and can't wait for the cultural zeitgeist among developers to swing back around to valuing users as human beings and living up to the trust they place in us.


Apparently, there's a term for the general practice, it's called "link shimming" - https://www.usenix.org/conference/usenixsecurity20/presentat....

I had some fun in my grad days building a private information retrieval (PIR) scheme so the shim server can do its job without knowing which link you clicked: https://github.com/pncnmnp/shimmey


i tried visiting kagi now, and there is no way to search if you dont signup first?

also this page doesnt have a submit button and pressing enter doesnt submit the form either:

https://kagi.com/html


The right click copy thing annoys me to no end.

I switched to DuckDuckGo around 2018 or so, which was when I realized that Google's results quality has deteriorated so much that I won't lose anything of value by doing so. This thread is how I learn about the atrocities that Google Search is committing these days.

Which is an elaborate way to say: you don't have to live like this.


I don't. At this point I just ask the oracle in natural language. There's so many to choose from. Why spend hours wading through only vaguely related things? You can ask it for a direct answer or you can ask it to teach you about the topic.

Sometimes there is joy in discovering things for yourself. Also, the oracle currently relies on humans publishing new information, what will you do when the incentive is gone and the oracle has nothing to harvest?

> Sometimes there is joy in discovering things for yourself.

Agreed. And webrings still exist in certain corners of the internet.

> the oracle currently relies on humans publishing new information

New? When I ask it how to do something with a piece of software it's perfectly capable of answering me by means of direct interrogation of the source code.

I'll grant you that it owes ~all of its knowledge of the world at large to having been bootstrapped using the more or less complete body of humanity's published works. However I think that was merely the quickest path to bootstrap it as opposed to the only one.


> This may seem like a small thing, but a series of misteps and chronic failure to adequately advocate for users is what has led us to the toxic waste dump that so much of the Internet has become today.

I think it is both monopolization and cheap money. Something in economy favors big companies and disfavors competition.

Cheap money means that a selected company can run at loss, kill competition and then enshittify to start earning. By that time it is too late and too easy to buy or destroy smaller competing company. And you get access to funding by being charming to VC, by being the kind of sociopath they like to see.

Capitalism works when there is a competition. Not when there is an oligarchy.


I was suspicious when they started obfuscating URLs in their own browser, then on their SERPs, and now this...

For many years, I had my filtering proxy rewrite the URLs in the way mentioned in the article.

Almost exactly a year ago, Google stopped working without JS. I stopped using Google.

Now they're upping the game, and as the article (which is a bit of marketing itself) admits, those who have the resources can still blast through these obstacles while those who don't are locked out.

Since the article brings up "AI scrapers", I'll just point it out as being the latest scare-tactic for coercing people to give up the privacy, anonymity, and (browser) freedom of an open interoperable Internet.


Direct URLs in Google search results have been replaced with redirect URLs in the form of www.google.com/goto?url=<opaque base64 string>.

The base64 data appears to consist of a very basic protobuf structure, containing a long string of bytes in field 2 which presumably identify the URL.

Sometimes, these redirect URLs take a perceivable amount of time to load, which is very irritating.


> Sometimes, these redirect URLs take a perceivable amount of time to load, which is very irritating.

Great. On top of my on-going battle with Windows + Firefox + DNS/TLS resolution sometimes stalling for seconds at a time, another few second server-side stall is introduced.

I swear that every day modern computing scenarios get slower and slower instead of snappier and snappier.


Wow - this exact same bug has been happening to me too. I gave up on troubleshooting it after the first few attempts came up with nothing, assumed it was just unique to me.

I had no idea others hit this, I just assumed it was some wonky setup I have locally. I hope we both figure it out one day haha (or Firefox does)

I searched for this a few months ago and found some mentions of this bug, but yes it affects me too. Can stall up to 20 seconds+ sometimes. Chrome is fine.

> On top of my on-going battle with Windows + Firefox + DNS/TLS resolution sometimes stalling for seconds at a time,

Have you tried disabling HTTP/2?


> Sometimes, these redirect URLs take a perceivable amount of time to load, which is very irritating

I do a bunch of work in remote areas with low-ping/low-bandwidth networks. The “hold on a skimmer while we round trip to Google” dark pattern makes the service unusable for me when I’m on-location.


The biggest problem with this is if the target URL doesn't load but also doesn't quickly error out, like e.g. many .gov sites in Europe (seems like they are just dropping traffic from non-US IPs).

Now you can't load the page and can't easily (using only the browser UI) get a link to paste into archive.org or archive.is to read the page.


When does that happen? I still get a normal address:

https://www.google.com/search?client=firefox-b-1-m&q=direct%...


Try it in a private browsing window - if you are logged in, Google will link directly to the result, but if you aren't logged in, it appears to be redirecting through their /goto endpoint

Are the bytes the raw url, or encoded?

The link you followed when you clicked hasn't a direct link for years, decade afaik (they mangle so they can see what's followed). The page used to show the direct on the search text but now it shows some stand in for it - sometimes. You can see the direct link on the bottom of the screen when you hover - sometimes (and sometimes you see a mangled link). Sometimes the google link contains the original link in the center also[1].

The situation seems to vary from result to result even on the same page of the same search - at least on the test search I just did. You can figure out what happening to an extent but this very inconsistency seems to speak to a dystopian quality to today's information gatekeepers.

[1] Example. https://www.google.com/url?sa=t&source=web&rct=j&opi=8997844...


I've just checked this again using a google account where search result pages are still following the old behavior: It looks like the 'href' attribute is the direct link, and the 'ping' attribute is the /url redirect link you are referring to. So it looks like it is actually sending me to the direct link, it just also requests /url at the same time in order to log the click. This means the user was not waiting for the logging/redirect request to come back.

Which is exactly the game theory that was predicted when some browsers started ignoring <a ping> to "protect privacy". If your browser supports ping you get ping, otherwise the website gets the data anyway but with a worse user experience.

While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be to build a set of sites of most interest to you and then prioritize those in crawling.

While typing this out I remembered https://en.wikipedia.org/wiki/Google_Search_Appliance which I never personally used but shows feasibility for the idea. I'm pretty sure one of the newly-announced Macbooks is more than up to the task of matching GSA's offering.


Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.

It would be interesting to see a decentralised, residential collective that builds and publishes an index. There are surely enough interested people on HN alone that would be willing to run software at home to scrape a small slice of the internet.

Something like YaCy?

Just like that! I will give it a try

The majority of websites try to do that but they do not catch as much traffic as they think they do. A starting point for a scraper is to run it on your home connection in an undetectable web driver framework such as zendriver.

Most of the stuff should be bypassable with browser automation but you'd need more compute to run a full browser versus a basic uri fetch.

I imagine it'd take quite a bit of shape for the index and be hard to keep it up to date unless you restrict what it indexes.


This was a knee jerk response to the first paragraph. They weren't talking about a general crawler, but a subset of Wikipedia, stack overflow, programming docs and github. You can download archives of all of those except github, and github could be queried using the api or GH cli

For some of that, you don't need to crawl. Wikipedia offers database dumps which you can download in one go. Lots of programming docs are managed in repos, so you can clone the repo instead. Even stackoverflow seems to have a snapshot dump (https://archive.org/details/stackexchange).

SearXNG configured as in the OpenWebUI docs is pretty cool. My "Hello World" with a new agent framework is teaching it to use SearXNG. Hook this in as a tool and the agent can answer a lot of questions.

SearXNG is more of a metasearch, the dude who wrote it pops in on here and is working on a cool sounding project that is more like a local personal search engine, I forget the name, but I've been meaning to check it out.

There is also Common Crawl.

https://docs.openwebui.com/features/chat-conversations/web-s...

https://github.com/brian-learns/xng-agent


SearXNG has been working pretty well for me. I had an agent write the MCP then do a couple passes comparing to server side LLM web tools and exa and tweaking and it works pretty well. I also added scrapling for fetch which covers pretty much everything but sometimes is a bit context heavy.

He's working on Hister now. I really like it.

Wayback Machine full archive is like less than 50PB. Let's say you could strip multimedia and remove every patterned data to compress that into about a petabyte. The per-bit cheapest disk right now is consumer Seagate 24TB(SI; 21.8TiB usable) at ~$500, or around $1200k for just the disks.

Doable if you had couple million dollars to burn. Cheaper than private jets new.


Seems like the most practical path is some kind of distributed peer to peer contraption where you can allocate some storage and optionally participate in crawling.

I imagine if set some constraints you could get index size down quite a bit but I still suspect it'd be hard keeping up with content churn.

Edit: Looks like enwiki bz2 is coming in around 46Gi which isn't too bad considering the amount of content it contains.



It's still kind of an open problem. There are partial solutions, but not yet really an integrated one. There is Hister: https://hister.org/ that builds some sort of drive-by index of what you're browsing anyway. And there are "true p2p" solutions like YaCy: https://yacy.net/ but it takes ages to crawl the open web (and lots of storage). There's still a lot of optimization to do in this space.

Until 10 years ago, i used Dash for that. It’s still around https://kapeli.com/dash

It’s instant, works offline, auto-updates, and includes all the websites you listed, and allows for custom ones too.


What do you use now?


Plausible. Figure 100 GB each of search index for Stack Overflow, Wikipedia, and GitHub issues, then add a dozen more for docs of all your favorite techs. So maybe half a terabyte. Download and build updated dumps of those once every week or two, and it'd work pretty well. Impractical, but possible.

What does this mean?

This text appears on this page https://www.autom.dev/blog/google-search-goto-links

> The real URL is in the Location header on /goto. Request that URL. Do not follow the redirect.

And this text appears on this page https://www.autom.dev/blog/google-goto-url-fix

> Do not follow the redirect. Read Location.

That's what a redirect is, reading the value of the location header and then requesting it. How do you not follow the redirect by reading the location header? Once you've made the request to the /goto url, with GET or HEAD, to get the location header, google knows you're interested in whatever it is putting in the location header and can assume you're going to go there, if you're letting the User Agent (curl or the browser) go there for you or not.


When Google stopped paid API search a few months ago, I looked for an alternative for my agents that I felt would be sustainable (one-time setup, then out of my mind). I quickly excluded SERP as I feard Google would pull exactly this type of shenanigans to cut them off.

I somehow found Mojeek and settled on it. I had never heard of them. Unlike Kagi, their business model is ads (so they hold no particular moral high ground). But they have a cheap, working paid API.

What I was astonished by is the quality of the results. For my uses, it's undistinguishable from Google. The conventional wisdom is that web search is a Google-sized problem. How did those obscure Brits pull it off?


I’m guessing that the knowledge of the techniques to support large-scale web search have diffused out of Google - it has been a few decades after all. Not to mention that the distributed system knowledge that used to live only in Google was either published by Google or cloned in other projects, usually made by ex-Googlers. And you can now rent capacity at scales that 20 years ago required Google to build lots of their own data centers.

Also, there is now a use case for paid search APIs - LLMs and agents - that essentially didn’t exist a few years ago. Not sure why Google hasn’t leaned more into this, but probably a combination of Gemini and Ads interests have combined to view their search index as an increasingly valuable asset, when if anything it might be corrupted already by these interests and therefore be less valuable.


I've been using DuckDuckGo for years now, ever since it became noticeable that two different people searching for the same search term would get two different results back from Google. Meaning they were no longer completely reliable: they might show one person a result that they hide from the other person by burying it on page 3 where few people ever look.

DDG's search results have been poorer recently than they used to — I often see completely unrelated results (to the point of my saying "Why in the world did that come back as a search result??!?") starting from page 2. And yet, I still use them, simply because they aren't Google.


Speaking of page 3, Google's been asking me to sign in if I try to go beyond page 2 of search results. What a wonderful modern WWW we live in.

That somehow hurts me to hear more than anything. I remember days spent in my youth trawling through Google search for anime fandoms and homework answers. Joke used to be that beyond page 1 is the definition of desperation and in the 00s that was definitely true but I discovered a lot of cool distractions in that desperation. This feels like they killed Google Reader again.

Don't know how modern this is, but I remember being disappointed that despite Google claiming it had billions of results, you could only see a few pages' worth.

Looking now, I've just noticed they no longer give a count of results.


I don't know how global consistency is actually useful, but I can easily think of ways it is less convenient.

For example, DST starts and ends on different dates in the UK and the US. Which dates should google return when someone from either country searches for just "DST end date?" Someone lives in Orange County and searches for "Orange County Sheriffs," which one of the eight Orange Counties should google return?

These are both examples where a localized — not even customized, just localized based on IP addresses — search results will easily help reduce headaches.


I wouldn't mind that one. That a man in the US and a woman in Japan would get different results for a search for "sushi restaurant" is perfectly reasonable (even if the woman in Japan was searching in English).

It's when two people in the same neighborhood got different results for the same search that I said "wait a minute, they're personalizing search results now for ad-targeting purposes" and ditched them. The potential for them to deliberately hide things from you was too great.


I understand the potential for abuse, but it is useful when I search for python and get programming languages and someone else gets snakes.

When I search Python I should get the programming language but when my neighbour searches Python he should get snakes. That is a valid use of personalised search.

Good example, which is why both you and the other commenter gave it. And if that's as far as it went it wouldn't have bothered me so much; it was when conservative people I know were getting different results from liberal people for the same search terms that I started to realize the potential for an echo chamber, and decided it was a bad idea.

It's good for your mental health to be exposed to arguments you disagree with. Many times they will be bad arguments and you can dismiss them, but it's good to know what the people who disagree with you actually believe. And sometimes, just sometimes, you might see something that makes you realize that it's your own worldview that was mistaken, and adjust your worldview to better fit reality. If your search results only show you echo-chamber results that already agree with yours, you'll rarely learn the areas where you were mistaken. (And nearly everyone is mistaken about some things; the person who is right about everything is a rare creature indeed, and you should never gamble that you're that person).


If the query is just 'python', then both you and your neighbour should get some links for the language and some for the animal.

If it only gives you the one you already know about, it's a useless tool.


I’ve been using DDG for a long time now. Results are getting more and more spammy. Way too many AI-generated / scraped websites nowadays.

Try Brave.

Do they still mitm sites to put their brave crypto donation buttons on, so you think a content creator is getting the money?

The age of internet search is over. The age of Cloudflare has begun. It wouldn't be possible to build a search engine now if they wanted to... and there wouldn't be anything to search for anyway. The non-corporate internet withered into dust and blew away in the wind.

If you could find what you want, how would they ever sell you what they want you to buy? And I'm not just talking merchandise, though that too. Your political narratives, your values, opinions, everything. And everyone likes it so much they just sit there scrolling and swiping and tapping.


I'm hoping for a future where we create static HTML pages again styled with a bit of handmade CSS, because we're so tired of bot attacks and long loading times. Then suddenly a Cloudflare network becomes absolet.

people have been putting their fuckass blogs with 0.5 visitors a day behind Cloudflare long before the increase in bot traffic. and that increase makes fuck all any difference for them anyway.

Cloudflare is quite easy to bypass.

Do you have a source for this claim? Cloudflare gives me endless captchas on almost any website.

Yes, I bought a cheap residential proxy service and used it to scrape a site.

Ironic that I'd need something like that to stop getting blocked from a residential IP

It's probably cheaper to switch your ISP to one that your target site does not block, than to pay someone else to buy that connection and proxy all your traffic through them.

Sorry, I don't consider that a bypass, or easy.

I think curated directories from prehistory might make a comeback eventually, maybe combined with a search engine that indexes only the whitelisted websites. ironically, we can use LLMs to filter out AI slop by measuring the signal-to-noise ratio, which is atrocious for AI slop and ESL slop that preceded it.

2. They are MS Bing, how is that better?

DuckDuckGo is owned by Microsoft? Are you sure? You're not confusing them with Bing, by any chance?

No, sorry, it uses MS Bing as their main source, that also explains the low(er) quality

Personalized search results are not ipso facto unreliable.

It means they're capable of burying news stories that would contradict your worldview and pushing news stories that support your pre-existing biases, leading to more engagement from you (a win from their point of view) but also burying you in an echo chamber. And unless you were in the habit of doing the occasional search in Incognito Mode, you wouldn't know. (And even then, they probably would be able to put together enough clues to figure out your identity even without your Google login cookie).

Yes, the fact that they could do it does not prove that they were doing it, not right away. I ditched them as soon as I found out that they could do it, because I was absolutely certain that eventually, they would end up doing it. And I wanted neutral search results, not biased ones, even ones biased towards my own point of view.


I subscribe to The New York Times, if I am searching for a news event, I’d appreciate a site that’s not paywalled and I trust be the first result if it is reasonable.

How u reliable personalised search results are depends on who controls the personalization and for what purpose.

I would love a search that lets me ban results from certain sites. I don't like search that's showing skewed results for reasons I don't know or control.


Anything that isn't repeatable is by definition unreliable.

Random number generators notwithstanding.

Same seed gives the same sequence. Use that to debug Monte Carlo simulations all the time.

The definition of "not repeatable" or unreliable in an rng is if it was random sometimes, and predictable other times.

They’re unreliably unreliable.

"Combined with earlier moves like removing &num=100 and tightening BotGuard/SearchGuard, Google is steadily raising the cost of naive SERP scraping."

Another "move" is suing companies like Autom, e.g., SerpApi

Google's Amended Complaint from their suit against SerpApi

https://ia801008.us.archive.org/25/items/gov.uscourts.cand.4...

"30. Copyright holders have authorized Google to implement access controls like SearchGuard for the content they license to Google, and in some cases insisted that Google do so. Googles authorization takes many forms. For example, Google has an agreement with a prominent licensing partner that holds copyrights to millions of works that it licenses Google to use in its Search results. Under the parties agreement, versions of which date back to 2017, Google is not only authorized, it is obligated to use commercially reasonable efforts to safeguard the licensed content against unauthorized third-party access. Other license agreements contain similar obligations. For example, another major content provider requires that Google ensure the content it licenses will not be available for download by third parties, thereby authorizing the implementation of technical access controls."

"31. In other cases, Googles authorization to implement access control measures like SearchGuard is part and parcel of the grant of licenses themselves, as Google and its licensors recognize that the value of the licensed rights would be undermined if others were free to access, take and resell the licensed content without restriction. For example, Google has a licensing agreement with Reddit, under which Reddit licenses Google to use the copyrighted content of both Reddit and its users in Search Services."

"32. Googles licensing partners have also expressly requested that Google prevent unauthorized access to licensed content. For example, when Reddit suspected that scrapers like SerpApi were accessing, taking, and reselling the content that Reddit had licensed to Google, it specifically asked Google to employ technical measures to prevent such unauthorized appropriation."

But this does not account for material that is not covered by the "license with a prominent licensing partner", its license with "another major content provider" or its agreement with Reddit

Google not only uses SearchGuard on SERPs containing links to the content covered by these licenses, it uses SearchGuard on _all_ SERPs

Google needs more than a "goto" update. It needs to update its terms to require _all_ copyright holders for the materials it has indexed and cached to give Google authorisation to use "technological protection measures" to deny access to certain members of the public, e.g., Google's perceived competitors including any Google user who "searches too fast"


They are allowed to scrape everybody else but get their feelings hurt when they get scraped....oh yea they respect robots.txt. Guess what, scraping everything that is public is legal.

At this point, we can all just drop SEO [0]. We are writing content for a robot that hides the source of information.

[0]: https://news.ycombinator.com/item?id=49665572


We do GEO now and it's not to make the robot link to us, it's to manipulate what the robot thinks.

How do you even measure GEO, sometimes I try searching for my app and sometimes claude recommends my competitor that has 10x less reviews and is objectively worse saying is the most popular.

To measure it, you could ask it 100 times and see how many times it returns each answer.

Wait, so what about the keywords in my meta tags?

but you can click the link and reveal the information. what is being hidden? from who?

Like the OP says, it's being hidden form scrapers. Like you I don't see a reason for normal users to be concerned.

Unless I want to see what the link is before clicking. Or copy the link.

Or navigate to the link without waiting for however long it takes for Google's redirect endpoint to respond.

Why are you defending the ruination of one of the basic principles of hypertext?

Why are Google apologists flagging my other comments?


Because Google search isn't a hypertext document, it's a web app that presents dynamic ephemeral content to the user. There's no need for stable hyperlinks, since either you will click them within a few seconds, or they will disappear forever, and it shouldn't be a problem to indirect via Google when clicking them, since you just made a request to Google to get the link in the first place.

Ask the grifters and bots. If you don’t adapt get ready to get bulldozed.

Google seems to be swimming in hypocrisy, as they develop and use scrapers.

PS: I do not understand why people keep using Google search, the results turned biased and mediocre.


Hmm that's actually pretty clever. They can serve each result page slightly different encrypted links and it should be obvious right away if it's a SERP bot (trying to grab a page of links) or a human that just picks a few here or there.

I wonder if this would also work on other sites getting hammered with bots. Allow each anonymous user 1 "real" page load then turn the rest into encrypted links that the web server can decrypt. If a session cookie with reputation exists, stop screwing with the links.

Kind of annoying but it'd allow tracking if the same agent/bot is churning through IPs/User Agents.


So why are we angry about that ? I mean the end result for the users are exactly the same, it matters only for bots.

Google have such a (justified) bad reputation that whatever they do, people assume it’s entishification. I don’t believed it is on that matter.


It's rather frustrating that after they crawled my website to use for their AI without my consent, and using that AI to slowly strangle my traffic, they're not happy that _other people_ are doing the same to them.

This is what capitalism's all about. Take as much as possible and give as little as possible if you want to win. Ideally, give negative amounts.

As opposed to you communists who freely deal out death.

>”This is what capitalism's all about. Take as much as possible and give as little as possible if you want to win. Ideally, give negative amounts.”

Your description is vague enough to encompass every system that I’m aware of, from feudalism, through to mercantilism, communism, and capitalism. It may best fit communism and feudalism, both of which are most notable for fostering negative-value-added firms.


I do not understand the type of psychosis which causes people to say that capitalism is communism.

You do not understand the comment you're replying to either

But if someone says "dogs have four legs, and cats have four legs", they aren't saying that dogs are cats.

It breaks the social contract of the web. My user agent should be able to tell me where a hyperlink goes without first clicking on it, and perhaps act differently based on that information.

Now, all Google search results appear to go back to Google.


TBF this is more a browser problem than a website problem to my mind. Redirects shouldn't be handled in such a cavalier manner. I should be able to configure the browser to stop and confirm the destination any time the domain changes without direct user interaction.

1. It adds extra latency due to an extra hop.

2. It is non-bypassable tracking of every single click.

3. It does not allow you to inspect the URL to see where it goes to before visiting.

4. It breaks the feedback signal of extensions which redirect sites. E.g. Fandom wikis are shit so I have them redirected to equivalent much better wikis like. But now any such redirection has to be done post-click tracking meaning Google still believes I want to see the Fandom site.

5. It breaks extensions which hide certain shit search results based on URL.


Privacy-wise, it enables them to track the result you click on.

Also, though more niche, it would make archived search result pages (e.g. on the Wayback Machine) less useful.


I think they always knew what link you clicked on (when most of the web has a ga4 tag on the other side of the link)

You can block Google Analytics with your browser. You can't block needing to ask Google for the URL of the result you want to visit.

If you're that concerned, use another search engine

I have to anyway because google always gives me endless captchas, as my ISP frequently rotates IPs with people who can't behave.

That's just Google not caring about your ISP because you live in the global south. Everyone gets this. US citizens get the exact same problem but Google whitelists their networks.

Not just global south, happened all the time to me after first search page of the day until i stopped using them, northern eu.

> because you live in the global south

I don't, but thanks for confidently assuming.


At least Google thinks you do. And Google kind of has the power to decide where the global south is, at least as far as it pertains to captchas.

Aside from the list of current concerns that others have covered well...Google is very good at "boiling frogs". That is, rolling out unpopular things in phases.

This could be, for example, step one in the return of AMP, but with a new twist. Where google conveniently returns the content of your website, without directly sending the user to your website. With whatever changes it chooses to make.


For me it's annoying that I cannot right-click the link and copy the actual target. They are breaking the web.

Personally, I put a lot of effort into writing my blog posts, some of which have been really well viewed by humans. While mine is effectively a "hobbyist" site, it's nice when readers look at other blog posts as a result of reading the original one. Google's approach ensures that readers get a very narrow view of my content, so everybody expect Google loses out.

Aside from the privacy concerns and inconveniences this brings to Google users, it matters for users (of services) that rely on SerpApi too. Kagi for example has started showing these goto URLs in search results.

Well, that's rather the point, right? Kagi and other micro search engines sell a product that just resells Google search results while telling people it's a premium product better than Google. Obviously Google isn't pleased about that.

Yeah. I might be wrong but I think they only served direct links for a relatively short time in their history, early on in the 90's and in recent years with ping, which they used to track clicks anyway. At least half of their history they used either the 302 redirects or onmousedown link rewriting (which was terrible). And I'm not even starting on AMP.

Well, for one it breaks right click and copy link location.

In this case they assume correctly that it is enshitification. A search engine, that doesn't understand what hypertext and the web are supposed to be is shit. There are no good reasons to not have proper links. Whoever creates such websites is having other motives or doesn't know good web development. In the case of Google employees I have to assume the former. They are not acting in the best interest of the users, and therefore enshitify Google search.

Well, I hope ultimately this will lead to even more user loss for them. Fortunately, I myself don't have to suffer due to it, because I degooglified my life.


It's not exactly the same: when in hover over a link, I want to know where it's taking me. I want to know if the URL is pointing to a safe domain.

Google being the cesspool it is, can serve whichever (ad laden) domain that fits their own interests the best.

This is yet another move to remove power from the users.


It's sad that instead of searching things other people put up, we're basically asking sam or dario oracle to tell us the truth. The people should be furious. But we've internalized this idea that they are somehow better.

Not better, just much _much_ more convenient. And it doesn't have to be a frontier lab. I'm happy to ask any oracle that returns sufficiently good results which is an ever increasing number of them.

If you were scraping only Google with all the IPs you can get, then this change really slows you down.

If you're trying to fight scrapers on a small site, delay links can only flatten bursts. If bots can only scrape at human speed per IP, they can just scrape 100x as many sites at the same time. Once every bot operator does that, total traffic will return to the original level.


Can someone explain why this matters? Not being flippant I just don’t understand why this would be important.

It's primarily relevant because it makes scraping search results much more expensive, solidifying Google's effective monopoly on Internet search.

Google has previously tried to prevent scraping of search results using legal means, but courts correctly think that scraping of Google's search results should be legal, just as Google's scraping of the whole Internet is legal. This is Google's reaction to that.


This is the best explanation. They’ve been doing the same in Google News. Each entry comes not with a URL to the source, but with a hash. To resolve it, you must send requests to Google’s servers. Anyone who wants to create a list of URLs of sources automatically can therefore be blocked by Google now on two levels rather than one - the search for a list of results, and identifying the source URL for each result.

In effect, they’re removing attribution from the content they quote from other people’s websites. It would be interesting to see if courts object to that. It is one thing to crawl other people’s websites and display snippets of their work as your search results when each result is properly and transparently attributed. But if the text is quoted and the source is not there alongside it in plaintext, replaced only by a vague promise that, if you ask, we may or may not tell you where this piece of content is from, that is a very different deal.


> It would be interesting to see if courts object to that.

It is also a measure of how enshittified and exploitative thing have become, that we look towards litigious copyright holders for assistance...


You can opt out from Google scraping you though? In theory you can opt out of anyone scraping you (if people were well behaved). Google should get to opt out of being scraped too.

> You can opt out from Google scraping you though?

You can't. Google will ignore robots.txt in some cases (e.g. "The REP isn't applicable to Google's crawlers that are controlled by users (for example, feed subscriptions), or crawlers that are used to increase user safety (for example, malware analysis)").

robots.txt is just a suggestion that Google loosely follows.

https://developers.google.com/crawling/docs/robots-txt/robot...


Well just as a regular user, I think it is pretty annoying because if I look up anything on Google while in Incognito Mode and hover over a search result, I can see that maybe the top result is maybe Wikipedia, or Instagram, or some other less-known website depending on what I'm searching for. Now, that's all very obfuscated because I don't actually know where I'm going to land for sure.

Tracking. Tied to use account, length of stay, how many reclixks, etc for marketing by/ads and surveillance.

But what’s stopping Google from doing this already?

I copy a tracking link from google in my reply here. You click on the link. Google now connects me and my search session to you and knows where I gave you this link from. If I get the original url, all google knows is that I copied that link and nothing else.

JavaScript being disabled - Google was already sending analytics pings when search result links were clicked on, using JS.

<a ping=""> works without JS in Chrome and Safari (but not Firefox) https://developer.mozilla.org/en-US/docs/Web/API/HTMLAnchorE...

JavaScript is required to use Google search at all, no?

Oh that is true, I forgot they changed that.

It doesn't work without JS though.

Pings are blockable in Firefox and so are the event analytics requests triggered by Google's on-page JavaScript.

They were already tracking everything you click but for example if you want to send a link to someone you can't copy the link from the Google result and send it to them, you'd either send them the Google tracking link or go to the website yourself.

If you make money trying to resell Google's free services and data, this is a disaster.

As a regular user, it's a minor annoyance.


it sounds like it primarily matters if you are a customer of this company, one that is building a search index off of urls scrapably hardcoded (or at least so as to be easily unencodable in non-realtime, it sounds like?) inside google search result redirects. in theory, there could be noticeable consumer user impact, but ... it would have to be a pretty large theory

It's to make it more difficult for scrapers. Google have been able to do all the tracking they've wanted for decades.

Think about it: if you want to set up something like jina.ai you need to build your own index or piggyback on Google. My guess is they use residential proxies to fire requests to Google then scrape the URLs.

Opaque URLs now give Google another gateway to detect circumvention of their anti-bot controls, helping them monetise their search index rather than allowing providers like jina.ai to succeed.


The security implications of this are very severe when you consider the amount of people who google government websites, banking, crypto and others. And google will happily serve you a phishing website either in ads or results.

Can you elaborate?

When you search Chase Bank and get a link that says Chase Bank you can't know if it's Chase Bank until you click on it

Sure. And why does that matter?

It should be obvious, they are doing this for a reason to benefit themselves.

As others have elaborated, the reasons are so they can track who you are and sell your profile advertising.


> the reasons are so they can track who you are and sell your profile advertising.

What? Like they weren’t doing this before? Obviously Google’s telemetry is tracking every link you click regardless; there’s no extra tracking benefit to this.

The reason they’re doing this seems to be to stop competitors from scraping their search results.


It benefits on the other end, if I share the link and you click on it, Google knows you got it from me and from where. People 100% don't click through to share original URLs.

Dont worry in a year or two, Google wont even redirect you to the actual true url, instead everything will be a page with all links rewritten so all http is tunnelled thru them.

Doesn't this just describe AMP?

what happened to that anyway? I remember seeing it everywhere, and it just... stopped?

False start.

If I hadn't already ditched Google long ago, this would be more than enough to chase me away.

They are forcing my further and further towards kagi. I do hope the fawning on here is at least partly justified.

Kagi is what I settled on a couple of years ago. I do think sometimes the fawning is over the top, but it is a solid search engine and tended to give me a bit better results than Google out of the box. The real win, though, is that you can give various sites a weight, so the search results will prefer or avoid sites according to your desires. Once I had that going, my search results tended to be much better than Google.

Yep, this. I was delighted they implemented this simple yet powerful idea. I don't even remember when I have needed to go to a second page of search results in Kagi and spammy clone sites which copy stack exchange verbatim are banned into oblivion where they belong. Fuck up once doing that shit, and you are forever out of my search results.

And then there is the very handy assistant, that I can jump into using follow-up questions to the quick AI queries one can choose to use or not to use by appending a "?". Again a simple idea, which empowers the user.


Yup, same here. Been using Kagi for years. It's boring, it just works. Hoping they can stay that way.

My ongoing concern with Kagi is they always seem to be focused on sidequests like their Orion browser and their LLM-powered Translate tool. Maybe that's interesting for some people, but I can't help but feel like I just want a damn search engine than works.

I also have no interest in their side projects. It doesn't bother me that they have them, though. As you say, as long as the search engine works and the price is acceptable, I'm happy.

Yeah, that's the thing for me: filtering out the SEO crap that Google happily serves up.

Google's actual search results are a waste of time visiting, both for the mindless CEO content and the ad-laden, analytics-happy, javascript-heavy websites.

So I tend to use the AI overview. But plugging myself into the all-seeing corporate oracle, that grew on all the web's content, and now seeks to supplant it seems unseemly.

It's just sad that kagi will likely only be a fringe thing, and google will continue to promote these foul, foul, mindless websites and then supplant them with its AI.


I don’t fawn, I just pay and use-and-forget. It’s almost to the point where it’s a little mental bump when I have to make a browser use it again, setting up something new.

Kagi has a pile of features I’m not getting the benefit of too, I’m sure, because it’s just-search, mostly, to me.

That’s fine; I know what I’m supporting, and I know that I’m the customer and not the product.


I’m not fawning over it but it’s

- usually better than Google

- has actually useful features (smush BS listicle articles together, weight/block sites)

- if you don’t use any searches that month, they’ll just straight up refund the month subscription automatically

- doesn’t annoy me constantly


Why? I don't really see why this is bad, other than if your company blocks url shorteners/redirects. What are you losing?

Because I wouldn't want Google to know what links I click on. When I used to use Google, and they added such redirects that included the destination URL as a parameter, I'd edit the link to make it just that URL before resolving it. This new scheme would make that impossible.

They own the JS on the page, they don't need the redirects to know where you clicked...

They already knew which links you clicked on though…

How? I never allowed Javascript or anything.

I use a Firefox extension to rewrite those Google redirects into plain links, to remove Google tracking of which links I click. The extension is broken now

Blocking URL shorteners and google-ad links yes! Personally for me it's also the fact that this is effectively an unresolvable URL shortener, store that link somewhere and it will most likely be dead. Can't copy link anymore and paste it on a notepad or chat app to check it out later as there isn't a guarantee it will load at all (ie: the problem with url shorteners).

> The real URL is in the Location header on /goto. Request that URL. Do not follow the redirect.

What does this mean? Isn’t the location header the redirect? Am I not following the redirect by requesting the location header url?


Can confirm that when using google not logged in. Now when sharing a link from google search, I won't get the actual link. This will certainly help google's tracking.

Google has gone so bad over the past few years. You only get like 8 results per page. I remember there was a time that I wonder how a site get reached if it ranked on the second page, when I can set the number of results to be 50. The censorship is also really bad, and google doesn't even tell you the results are censored, returning totally nonsense results while other search engines work normally.

I've been supporting Brave search which returns 20 results per page and has other features. It used to be not good a few years ago, but now the results are often better than google's.


That's it. I'm done with Google search. I just found you can add Kagi to Safari: https://apps.apple.com/us/app/kagi-for-safari/id1622835804

This is such an incredibly annoying and deliberate defect. I want an extension that will let me resolve the true url without visiting the site, or even rewrites the entire page to show the url.

1789154298 | Google's Emissions Climbed 48% Since 2019 due to AI | https://www.gadgetreview.com/googles-emissions-climbed-48-si... | https://news.ycombinator.com/item?id=49663855 | 0 comments

I've never heard of this particular SERP provider, but some marketer is very excited they wrote this blog post right now (500+ upvotes on an seo blog post).

And their fix here - they just resolve all the urls- which I suppose could make the service slightly more expensive? But otherwise isn't that what every similar provider/ crawler etc will do and this change will only hurt users?


At least they seem still provide results for my searxng instance. I mean sure, they are horrible but duckduckgo just blocks most queries (and I'm the only person using the ip / seraxng instance)...

Next i'll do is to implement tavilly, exa, tinyfish etc. as search engines for searxng. No agents, no mcp, just their search api endpoint.


Brave search (free) or Kagi (paid) are able to replace Google and not feel like I'm missing out.

Brave has its own independent index which is cool.


I use kagi exclusively because it works. I get the results I need every time, and it's not annoying. I haven't used Google in years. Simply no need. I really hope people keep paying kagi so they stick around.

I hope you also know they're using Yandex as one of their indexes.

https://kagifeedback.org/d/5445-reconsider-yandex-integratio...

Uruky is the way. 2x cheaper also.


+1 for Kagi, even if Maps and Shopping and Images aren’t as good as Google (tbh their Images search is good enough 8/10 times) they’ve got the actual web search stuff working quite well.

> Combined with earlier moves like removing &num=100

Removing this made google search horrible to use. I often use command+f to quickly identify relevant search results, but doing it on 10 results at a time is so laborious that I just don't bother using Google search, resulting in less searches and use of other tools instead.


would you mind sharing such other tool?

I refer to tools like LLMs, unfortunately (not search engines). If anyone knows of a search engine that returns 100 results, I'm very keen to learn of it

I had tried now but I don't see the goto. Is it possible that has been deployed only for USA users ? (I work and live in UE). If so: probably a vpn can help you for a while. In the future: I think I'll really go for payed search engine.

I... Don't see it? It's the result page right? I just search some random string on Google and the results are all direct URLs. Do they get resolved via javascript after page load and replaced automatically? Or am I looking at something else?

What I see now, and it's been like this for a while, is this:

You get the results and they do have direct URLs. But then, if you do some things with the link, e.g. right click to open it in a new tab, it swaps the URL to the indirect one. The idea is that initially you see a normal link, with a normal URL which will be displayed correctly when you hover the mouse over it, but right before you click it, it's swapped for the indirect one.

So, they have been doing stuff like this for a while and it has been somewhat fluid, because the swapping can occur on different events and I have also seen it load with all the links pre-swapped to the indirect ones, sometimes.

So, yes, what you see may be different and you may get the indirect URLs swapped at different stages.


I was also super confused because it doesn't do this if you are logged in (to google).

I ran the same search in an incognito window and it showed the /goto links.


Thanks! This was it. Saw the /goto links in private window.

Are you logged in to a Google account?

It looks like it isn't yet rolled out to some users, try with a different browser session.

Is it a move to sell more of the paid Google Search API calls? If so, that's a sign of distress.

Surely, the cost of serving search results to bots isn't that high


It's funny. Recently I looked into what Google is doing (see at https://blog.miloslavhomer.cz/how-google-sees-your-site/).

It's a lot of work to get the data to build an index. Why would they give it to everyone for free?


Becuse they are a monopolist and we can require certain things from monopolists.

Similarly, Bell Labs was kind of required to release transistor for anyone to license - it was a part of social contract that they were allowed to maintain their monopoly in exchange for releasing certain parts of technology.

Alternatively, they could be split up and their indexing division made an independent company selling to anyone on a free market.


I tend to agree - businesses should give back. But they can also keep some secrets.

R&D has steep cost. Publishing it all gives your competitors an advantage.


It has steep costs, but with monopolists they have other ways of extracting value from the inventions.

Awesome book about the history if Bell Labs - virtually all semiconductor tech we use today was created there (transistors, ics, solar, lasers, fiber optics, telecom satellites…), and they had to license it to be allowed to maintain their monopoly status.

https://www.amazon.pl/Idea-Factory-Great-American-Innovation...


Thanks for the recommendation, I do see your points. We should be treating monopolies differently.

Thanks :)

Another cool book about back and forth between monopolies and decentralisation in our space: https://www.amazon.com/Master-Switch-Rise-Information-Empire...

It also has an audiobook.


I really like the idea of turning the search index into a public utility. It is one of those natural monopoly coordination problem things. Just quasi nationalise it for economic efficiency. Ofc google can still sell adds against their own ui (like everyone else). Hopefully this move will move that idea closer to reality.

With the current administration? They would manipulate the results to ridiculous extent

So would the previous and the next one, this is not situational.

What are you going to steal from Google?

The Internet has been dead for years and, after the Scrapocalypse, the small living remnants are behind a login wall.

Google can not provide you anything you couldn't find on either your local ZIM archive or the Media you consume.


Is there any market for an open-source internet search engine that is paid for by honest ads?

Been using Brave Search now for a while including its Brave AI and aside of sporadic times I never needed Google (albeit Brave Search is slower to Google, you get used to it)

The udm=web parameter has been doing the job for a while now. It's a shame you have to know about an undocumented flag to get the old behavior.

That flag is the result of clicking on the "Web" tab of the results page, so it kinda is a documented feature. All the tabs have their own udm value (e.g. image search is 2, AI mode is 50, etc.).

I use ChatGPT for almost all my searches now. I’m not joking.

Whenever I've tried this, it seems okay if you need an answer to a question, but plain bad if I'm looking for a specific page.

e.g. I'm just now looking for the menu for a local restaurant. "restaurantname menu" in Kagi (Google would presumably be similar) returns a link to the menu as the first result in about a second. Or "restaurantname menu !" goes directly to the menu in about a second.

Meanwhile, searching "restaurantname menu" in chatgpt takes about 5 seconds to return an embedded map from mapbox showing the location of the restaurant. If I click the restaurant pin on the map, there's no menu link, the 667 reviews have no link or way to view, and the restaurant description literally says "I don't have enough information to identify which local business <restaurantname> refers to."

Below the map there's some text: "If you mean <restaurantname> in <place>, here’s the current menu. <restaurantname>". The <restaurantname> link just opens the same card as clicking the pin on the map.

After that there's a bullet point list of the menu that ommits a ton of detail and options.

After that there's finally a link... that I can click to open up a popup at the bottom of the page with an actual link to the menu.

This was literally the first thing that popped into my head, I didn't have to put any effort into finding a query where chatgpt falls on its face.


You might have better luck using google maps (even the web version) to find menus.

Yeah, I would... but also a search engine works even better than a map if it's a restaurant I'm already familiar with and just want a link to restaurantwebsite.com/dinner-menu

I wish one of those free AIs made a search product already. I don't want my search results to be interspersed with text

I guess that the web chat can have a search skill to remove the prose and give only links, plus maybe an excerpt of each result


Z.ai has a search MCP server, it should be trivial to use that to build a basic search UI on top.

Yeah, its way faster.

Its not because chatgpt is so superior. Its just because google search is dogshit.

They work on killing the web as we knew it and I fear its kinda working.


Can’t sell Gemini if they were to make Search good.

Streaming services already adding in ads to “ad-free” tiers they’ve now named “premium”.

Quality of life on the internet has gotten shitty while Reality Classic stays mostly the same, though more expensive.


Hopefully this also helps against the rampant CTR manipulation that’s made some search results purely a measure of spend.

Of course, this move is user hostile because you don't see what you're being sent to.

You can see the domain printed on the page, right?

I had to check :) I use duckduckgo mostly.

They still display the domain, at least, yes. Wonder for how long.


This seems good. I'm not sure why I should be upset that Google is preventing abuse of its service.

I think it’s the other way around - why should you care about a monopolist’s interests?

The enshittification continues.

Nothing good ever comes from businesses desperately trying to protect their moats rather than making their products better so they don't need to.


Thank Prabhakar Raghavan for that. After destroying Yahoo’s search product he failed up and did the same at Google.

In reward for making Google Search materially worse in every way except ad revenue, he failed up again and was promoted to a cushy do-nothing role.

Everything wrong with the tech industry, embodied in a single person.


He wasn't the one who merged ads and search into the same org. The blame has to fall on Sundar, surely.

> promoted to a cushy do-nothing role

The people who this happens to are not perceived as successes. Everyone knows they are gentle firings.


Well one has to look a level up and ask who the people are who hire such people and why

Do you have any source outside of the Ed Zitron article (which is highly specualtive?)

If Google's ad revenue is going up he is keeping the shareholders happy.

it's metastasized at this point

If you don’t like this, just don’t use Google folks. There are alternatives, use them.

Bing has done that for years. Hated it.

Most people i see using ChatGPT for any search related work in everyday life.

Maybe Google is observing this?


Yeah seems pretty obvious to me that most people are not going to be using a search engine in 5 years. In the sense of searching for something and combing through the results to find the answer.

Divide et impera. People need to learn. Get together, do things together, create alternatives, use that, maintain position, do not let outsider saboteurs neuter the project. Problem solved.

> The url parameter uses a custom, Google-specific encoding.

So that will be figured out, making the whole exercise void. Google knows this. What is the real goal here?


"Figuring it out" doesn't help if it's encrypted with a key only Google holds or a reference to some database record that you don't have. In either case, the only way to resolve this is to ask Google, which means they can track it as a click (and rate limit etc.).

Imagine it being the ID of the database row holding the actual URL. There's no way to figure it out without access to that database.

I use Google maybe 10 times a month.

Why can't agents just click on the link and follow it? I don't get it

Append (for logged out users) to the title.

Note: Article published by Autom.dev which, from a quick read of their homepage, seems like it scrapes Google search results in violation of Google's terms of service and sells those results to customers via an API. That's just my quick read of it, though, so this could be wrong.

Note: Irrelevant.

The reported behavior exists, does it not?

Anecdata: I've observed this behavior for several weeks already as a regular user without a Google account and there are countless comments from regular users reporting the same behavior.


It’s very relevant because the author has a conflict in interest. They can wax poetic about “open internet” but really their business depend on it.

Kagi also uses this

It's relevant to know when the source of information is biased due to a conflict of interest. The source's business is apparently scraping Google (and Bing, etc).

> The real URL is in the Location header on /goto. Request that URL. Do not follow the redirect.

what? the Location header is the redirect, no?


Gemini does something similar

No longer ?

Google has been encoding the target url for years.


I'm confused. Google has done this for several years already. I have a Firefox extension installed that reverts it - it's several years old.

They changed how they do it, such that it's no longer possible for extensions to revert it.

My monthly reminder to use Kagi instead

I genuinely forget I’m not using Google until I come across articles like this


wow thanks for sharing

I've been using Brave search for a year now, because there is no way to turn off Google AI search.

I don't miss Google at all.

Goodbye you shit company.


Adding -ai to the query works fine. Or like the other responder said, use the web tab instead of the all tab.

> there is no way to turn off Google AI search

&udm=web works fine for me, you can use a browser shortcut or extension to enforce it


"google.com/goto considered harmful"

"There's nothing to explain. You're trying to kidnap what I've rightfully stolen."

Google wants to build up a private web here. We already know this from AMP before.

Also, the search results are total garbage now. Just give it a try and you see how useless the UI is, with AI results first, then tons of commercial go-to entries and only then a few links that are often also totally useless and irrelevant. Google optimised towards crap.

We really need to get rid of Google. Qwant results are now a bit better than before, but still not that great, and I hate that its default UI is a clone of Google. We need more alternatives. Let's get rid of Google once and for all - it has disappointed too many people now.


now is time to start the ai to spam/reverse the - google.com/goto for url resolution, it can't be that hard to reverse engineer how the hash works. It was doable for twitter and YouTube, it will be doable here too...

Who the hell uses Google still? Kagi is the way to go! Even though it’s paid it’s worth it.

Google Search is dying, and it is showing all the standard symptoms.

why is this such a bad thing? it's not really any different from using a uuid as a user facing key, which basically everyone does.

and trying to protect your moat isn't automatically a bad thing. they clearly feel it's helping competition, so they're closing a hole. competition is good doesn't mean help your competitors.

- coming from someone who's been using fastmail as my personal for ~10 years because i don't want my emails to be backprop fodder


I guess someone made a website which google crawled and adding a senf made uuid to it is like google trying to own it rather than just being a true search engine just having index to it.

ISPs can presumably correlate the Google query string with the request following the response to the goto and so make a search index? I guess they would charge too much.

Do any large ISPs use visit data to feed into a search index?


ISPs do not see query strings since Google uses HTTPS. ISPs can only see the domain and IP address you are connecting to.

Doh, of course.

Can they even see the domain? Assuming you're not using the ISP's DNS of course.

Unless you're using DNS over HTTPS they can see the unencrypted DNS traffic. There's also Encrypted Client Hello, but they can also see which IP you're connecting to.

Also if the domain responds on the bare IP or there's nothing else they've seen on that IP it's a fair assumption. This is without even wondering if their router sends telemetry

Most people use the ISP DNS, but even if they don't, the Domain to IP is public information.

There's even a reverse pointer query that can be made to get a domain to IP from a legacy ARPA TLD, although it isn't 100% robust.


It hasn’t provided direct URLs for decades? Not exactly new behaviour.

I’ve got something that will blow your mind. Google now has tracking analytics, for get this, your business’s phone number. Some “Adsense partner” convinced our web admin to install a little script which changes your phone number on your website so they track phone call enquires back to search engine leads / advertising spend.

Yeah no thanks, that was creepy as hell and had it rolled back ASAP. You’ve got to realise the power these tech companies hold over your business. Don’t show up in the search results, someone lists your business as closed in maps, a tracking phone number goes dead so they can’t call you, you might as well have shut up shop and ceased to exist.


Would it be different to you if a third company (not Google) provided that same phone number tracking service?

Would that company then sell that data back to Google? The consolidation of control of information under a single actor is certainly a factor, but it's not that simple, and it's not the only factor. You shouldn't be so eager to outsource as much of your business intelligence as possible.

Does whether or not they sell to Google matter?

What’s the alternative? Roll your own custom phone analytics system?


It matters. The alternative is to use a product that doesn't leak your business intelligence to third parties.

This is nothing - Google has had things like the ability to track in-store conversions based on Android/Gmaps tracking for years.

Some of the things TV companies do will shock you too.


I've been using the 'ClearURLs' FF addon for quite a while now, and it gives you the URLs back in the results.

https://docs.clearurls.xyz/


That wouldn’t work at all with this url format - the url can only be resolved by requesting the goto link.

I recommend reading more than the headline


> That wouldn’t work at all with this url format - the url can only be resolved by requesting the goto link.

What is "that" you say wouldn't work?

The extension can work fine by requesting every goto link upfront. As far as your google searches go, this is just as private as before.

I would recommend you avoid judging a comment because of a metric it neither said nor implied.


Previously the clear target URL of a search result was added to the result-link (or div?) as an additional data-tag (aria or whatever, can't remember), and thus was able to be used with a userscript or similar to replace the Googleified tracking redirect link. And it was also part of the googliefied tracking URL, either way you were able to see the clear target URL.

That no longer is possible, you must ask google.com for consent to visit the result link before seeing what the final target URL is.


But the extension doesn't need to use that exact mechanism, and the person that linked it didn't mention specific mechanisms. The purpose of the extension is putting the real URLs back, and it can still do that.

And it can still prevent google from knowing which search results you click on, even though OP didn't mention that feature.


You can read the extension source, and description. It is obvious that what it does is strip url parameters and decode well known google encodings.

I said it can use a different mechanism. Why would I want to read the code for more details on the current mechanism?

> But the extension doesn't need to use that exact mechanism

How do you figure? How else could it possibly work now?


I said that in my first comment.

You do a google search. The extension resolves every link on the page immediately via the goto urls. This doesn't leak any information to google because they obviously know which links they sent you. Now the links are resolved and you can copy and click them and get clean URLs, without sending any information about which ones you're copying or clicking.


Isn't that the same mechanism as the one you were saying it didn't need to use, though?

I decided to verify the article's claims, forgetting that I had ClearURLs installed, and ClearURLs does not handle this new format.

The article also suggests that it cannot be decoded.


The reason this change is so bad is that it encrypts the actual URL, specifically to break addons like that.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: