As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification
I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". It is surprising to see how much they have stripped from our view - long tail results, actual results for product reviews and not ad spam, no preference for 20 page recipe sites.
There are still illegal streaming sports and movie sites everywhere (who knew) and all other seedy corners of the internet that have been neatly erased by Google. It makes me nostalgic for that brief window of time when the web was truly uncontrolled, when page rank had meaning and you didn't know if your search would return 0 results or 4,000 pages, which you could actually browse.
(Cleaned / decoded: 'You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each with a short text from the page, featuring the search terms. Avoid any pages that do not contain the search terms. the search request: %s')
Steep at 300 searches per month or about 10 per day. If you search often, it is too expensive. If you rarely search, not worth 5 bucks.
They are plainly trying to push people to their $10 unlimited plan. I would have appreciated if they allowed, say, 600 searches per month for $5 or so.
I find that I rarely do one search. It’s usually a search, followed by looking at the first few links and probably opening them in tabs, then refining the search terms to get better results. It’s an iterative process. So the 10 per day on average would be exceeded quite quickly.
I too feel the yearly price of $108 for one person, $155 for two people and $216 for a family is expensive. Ok, maybe the family plan is cheaper if you can find five others to share it with.
One of the wildest things to me: Kagi actually warns you before renewals so you can cancel first if you want.
That's confidence in the product right there. I'd argue any subscription which won't do this is designed to exploit people who don't need or use a subscription.
You know how many things are asking for "~$10/mo". People don't have the purchasing power they used to, and everything has a subscription model. Firefox with DDG is good enough.
As they said, you wouldn't pay $10/month for search. This is why free pay-with-your-data services won. Nobody wants to pay for what they use, so companies extract value in other ways and you can't complain about that if you're not willing to pay what it costs.
> Nobody wants to pay for what they use, so companies extract value in other ways and you can't complain about that if you're not willing to pay what it costs.
I'm likely old school here, but I never liked "paying for what I use" in the online world because it required revealing some personally identifying information (i.e. to make payment). My attitude is almost certainly irrelevant these days, with the scope of online tracking, sharing of data, and the near necessity of disclosing one's identity to sites other than one's ISP, but it was an attitude that was established in my mind nearly 30 years ago.
(I also don't trust the pay for privacy model. First of all, you can never assume that paying for something implies your data won't be sold. That was even true before the public Internet. Then, even with statements about privacy, you cannot assume that will be true in the future. If the company is sold, management changes, or the company decides the need/want additional revenues, that will change. And that assumes they don't use weasel words about what privacy means, or the ultimate "we reserve the right to changes this agreement at any time".)
Kagi support zero knowledge search through a Firefox extension and Mullvad supports zero knowledge VPN by putting cash in an envelope with your account ID written on the outside.
I mean, there’s more privacy protection on payment data than there is for email address and browsing habits.
Plus most companies just push out payment handling to stripe. So most websites have little idea who actually transacted with them to begin with (email dependent).
Google kept increasing its profits since then, and it was one of the most successful businesses out there way before they changed the business model to become enshittified. The argument that their business model changed because people wouldn't pay for its services doesn't really stand.
And Google's ad revenue is likely more than that. Or, if it's not for every user, the ones who would opt out of ads are the ones who can pay. You're cannibalizing your highest value users.
True. I got this as my first LLM subscription (was a long time skeptic, less so now but still a bit), but now I have cursor too so maybe I can even downgrade. Thanks for the reminder. :)
My time, attention, and “brain” are valuable. Having aligned incentives with my search provider and control over the results it serves me through weighting is worth substantially more than I pay for unlimited.
Kagi is the stickiest digital subscription I have.
It's relatively rare as of late that I need to specifically find something on the internet, and DDG does not suffice, so I rarely exhaust the 300 Kagi searches.
Most of the time I need to find out something, and in that case, Claude / Gemini / Grok / Perplexity / you name it work better. Most of my queries are technical by nature, not interesting for data mining on me personally.
The 300 searches a month work for me because I only use Kagi when I think it will give better results than others. And most searches are trivial so any free engine will do (like you have the website name and want the url).
FWIW Kagi has worked hard on their pricing over the last few years and has been trying different models. They're not "big search" the price is what it is so it can exist as a business. 100% fine to not be a customer obviously, but then you forfeit your license to complain about Google spying on you and ruining the web. Like they're trying to solve the problem. IMHO speaking just for myself I feel a moral duty to support them.
I agree on all your points. I'm just saying, that the $5 tier is not a honest to good option. The $10 one is but they can make the $5 one genuine too by increasing the quota. I'm sure they have done their research and maybe increasing it will cannibalize the $10 tier.
Good workaround! It's just sad that to achieve the previous behavior we now need to burn significantly more compute, and in turn energy, and with far worse performance and an inverted UX.
The most frustrating part is that they have all of the data, and oodles of compute, available to surface the same very functional experience they originally offered, but they would prefer the image of being visionaries rather than the reality of being useful.
I think the "original experience", whatever it may have been (and I strongly suspect there's more nostalgia here than actual worse results) would be very quickly overrun by modern SEO methods / tactics.
Kagi can be "better" than Google (for some very narrow definitions of better) because they basically don't have to think about SEO at all. They're small enough that nobody does SEO against Kagi. If their algorithm is sufficiently different that typical Google tactics don't work on it, even if it's much more naive because of that, Kagi wins.
Thanks, I needed to add a ?aep=11&atvm=2&udm=50 for AI mode. Wow!!!!!
>You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each with a short text from the page, featuring the search terms. Avoid any pages that do not contain the search terms. Do not simulate fetching results. The URLs must be complete and not truncated. No markdown. The current date is 2026. The search request:
I'm using Kagi every day, and unfortunately, it is not true anymore. I think what happens is that the blanket banning of IP ranges because of LLMs, and flat out incompetence and laziness, affects them heavily. I get results, or just a few unrelated ones many times every day, while Google happily returns what I searched for. I think Kagi's index is shrinking, or they really made their algorithm worse.
I slowly shifted to use LLMs because of this in the past half years, making the situation probably worse for Kagi. They were the last usable search engines for years now. I still pay for it, but it's increasingly not that useful anymore.
Same here. Been paying for Kagi for probably a year or so, but I find that it hasn’t really been able to keep up with at least my personal expectations. It feels quite slow (which I don’t remember it always being), and the search result quality is not always the best. On top of that, the 10$/mo is a little steep for me, given that their assistant and maps are quite sub-par in my opinion. I would have preferred them to focus more on the search experience, and less on the other products. Or give me a cheaper plan for only unlimited search without the other products.
It definitely is. The extension-based redirect in Safari is also unreliable. Furthermore, many of my searches (and others', apparently) unexplainably defaulted to Groningen in the Netherlands, even though I'm nowhere near the place.
I ended up not renewing my annual subscription. Since then, I vibecoded a small userscript that replicates domain ranking for DuckDuckGo – plus DDG already lets you turn off ads natively.
I like Kagi, but is it me or have their results been less than great for the past 1-3mo? Maybe the places they're pulling their results from have had some changes?
> You are Google Search from 2004. Given a search request, provide 10 links to relevant pages, each with a short text from the page, featuring the search terms. Avoid any pages that do not contain the search terms. the search request: %s
I'd say that the ranking, and thus the relevance, is noticeably higher. If I know what I want to find, Kagi is good at surfacing exactly that, not something vaguely related.
Oh that shorthand is fun. Curious it works at all.
But yeah I'd switch to Kagi if I had any more forcing function. Maybe without LLM's but google continously getting worse, that'd been the actual thing for me
For Google, the business side of you paying them doesn't work out as obvious as it might seem.
Advertisers are sold the idea that 1000 clicks/impression is X dollars.
The average impression might be priced at 1000/X$.
But you are worth much more than that for two reasons.
- you are a person willing to pay 5$
- Google can sell a “collectivized” product. The same way that eg collective farmer crop insurance creates actual value.
and the third reason (as consequence of the second) they dont want people to think about: it gives them more opportunity to fudge the value by making it all very complicated.
I'm willing to pay time or money to avoid ads, not to be spied and harassed by companies making me offers to buy something that almost certainly I don't want to buy. Those companies should pay me a percent of what I make them save by not buying ads to serve me.
And eventually when I want to buy something I research it and buy it and nobody had to pay anything to Google.
> That's why tools like uBlock and YouTube morphe are the only answer, negotiating with terrorists never works.
This is silly. The economy doesn't work if no one is willing to pay for goods. The content you're consuming won't be made if theres no market value to create it in the first place. Pirating is not a solution to our data privacy problem.
The capitalist economy is based on the explicit idea that everyone will do everything they can to extract value from everyone else, and it'll come to some kind of equilibrium. If you do not do this but you allow everyone else to do this, you receive worse quality products for ridiculous prices and do nothing about it.
In other words, you owe me $1000 as a condition of reading this comment. Your options are to unread the comment, pay me $1000, or pay me nothing to incentivize me to charge a price you'll actually pay (likely $0.00 for comments but not for everything)
I don't understand your logic. Why would someone that never sees my ads be valuable to me? Why would I care? Everyone spends money on something, the fact you spend it on adblockers doesn't make you automatically more valuable to advertisers. I have never heard of this "advert paradox" concept you mention and I have worked in industry for a while now. All this is just conjecture on your part and doesn't really correspond to how it works.
If you have $20 to spend avoiding ads, you have $30 to spend on cool products so the advertiser is willing to spend up to $24 to make you not block ads about a $30 product you'll buy (that costs $5 to make)
> If you have $20 to spend avoiding ads, you have $30 to spend on cool products
Yes, but if you are so against ads that you spend $20 to not see ads (I am definitely in this group) and then see an ad for a "cool product" anyway, then that product is no longer cool. Any time something like this happens, I legitimately develop a grudge against that company and their products.
I might have bought that product before seeing the ad, but there is no chance of that after seeing the ad.
That does not make sense at all. Why would you care about people who want to avoid ads? The real data shows very little people actually do that. How much you spend on adblocker is irrelevant to your profile for a business. They will get you some other way if you ad block on digital.
Sad but true. I paid for Prime to be ad free, easily worth 3$/month but sure enough now part of their catalog has become "only available with ads" or ads in their FireTV is before you hit the prime app (and then again when you do). So now their whole product is crap and I'm looking for a new streaming stick solution [anyone?] Sounds like it's time for a newcomer/disruptor like Angel Studios to release one.
Natural next step of the ad lobby is obviously to pay the lobby to make it illegal to bypass the money making machine if they haven't done so yet. I don't doubt that capitalism still works, it just takes time to "boil the frogs" dead [enough] to stop them from buying such crappy products.
The only “other engine” that matters is really a thin paid SERP façade in front of a scraped Google result. In other words, Kagi pays for access to an API whose implementation is basically live scraping Google and throwing away the ads.
There are a few other indexes that Kagi’s aggregator mixes in (such as Bing Search) but those haven’t been contributing much in terms of meaningful results. The sheer size of Google’s index still dwarfs all the others, including Bing Search.
So yes, Kagi doesn’t scrape but they pay SERP providers who do.
Marginalia likely returns better results than Google, when it returns useful results. DDG/Bing likely returns average quality results which can be better than Google's slop results. It's not true that Google is the only useful one.
They pay scraping sites for results, including SERP API. When I pointed this out, last time Kagi was discussed on Hacker News, an employee of Kagi said that they're trying to build their own internal index, but he didn't provide details.
I’m pretty sure Kagi’s entire pitch is that they’re an alternate and original dataset. Theirs is the only search engine whose data isn’t based upon google or bing’s databases.
To anyone who didn’t find Kagi useful in the past, I strongly recommend you try it again, as they have made huge strides in usability and quality of surfaced data in just the past year…. Huge strides since even a few months ago too, with more adaptive curation, data labeling, and UI/UX
> I’m pretty sure Kagi’s entire pitch is that they’re an alternate and original dataset. Theirs is the only search engine whose data isn’t based upon google or bing’s databases.
It's not. In addition to their own indexes the results include calls to "all major search result providers worldwide"[1].
There are independent search engines that only use their own index, like Mojeek[1]. Kagi is explicitly something like a search engine aggregator.
Absolutely. Where google is shaping results "to please corpo-political masters", Yandex is obviously influenced by the Kreml
I use Kagi as my daily driver. But there are some kinds of queries where my interests and Yandex's interests happen to align. Mostly anything corporate America wouldn't like, like evil copyright-infringing websites
Yandex censors the results, it is required by law, so I do not understand what are you arguing against. To be specific, any URLs, which are blacklisted and banned in Russia, must be omitted from search results. Which includes BBC and other Western media and explains the difference in images because in the Google's results the images come from BBC and Voice of America.
Also, if you try to search for "download Chrome" (in Russian) then the first result in Yandex leads to a scammy website: https://ibb.co/HDRJGgZD The real link is the second one, but I remember a year ago or so there was no official link at all. You can also note that official link to Google has a grey text saying: "the owner of the resource violates Russian law" (Yandex is required to show this notification).
Yandex has also been caught "accidentally" removing the site of Ekaterina Duntsova who was planning to nominate for presidential election in 2024, and showed fake/phishing sites instead.
So, Yandex is only good for searching torrents/pirated movies (which you can watch on rutube) and nothing more.
Naivete and scorn of autism in one short sentence - you're very efficient.
Yandex.com is the international front end to yandex, and not a portal for 'a handful of western autists'.
But on Russia's granular propaganda and disinformation efforts, you ought to read Peter Pomerantsev's excellent "This is Not Propaganda" before making silly sweeping statements.
You may not even appreciate the full extent of it. Long ago now, when Google took off and started dominating because it was “the best results” people didn’t realize that even though they were the best results, they were not actually good results. To be the best, you just have to be notably better than your nearest competitor, you don’t actually have to be good or even decent, just noticeably better. Today Google really has no competitors besides all the derivatives it has maintained to maintain an illusion of competition, the most obvious examples being DickDuckGo and Bing.
It’s somewhat similar with AI today, only the competitive landscape has majorly shifted with the Chinese models, something that wasn’t supposed to happen. All the sudden you no longer have a closed system of American models that can all be contained and managed as good enough to make people believe are “the best”; they actually have to compete in a far more competitive arena with the tension of the best model being the most competent model, and the most competent model, by virtue of AI, will be the least restricted, the least censored, the least guardrailed, the least controlled to present a telescreen and Hollywood type fantasy world where the good guys always win and … wouldn’t you know it … the good guys is always us, the people controlled by a psychopathic, narcissistic cabal that also controls what AI will tell you, the false truth.
It's a shame there aren't any good Chinese search engines. From what I've read the Chinese internet is mostly centralised around a few apps like WeChat.
Which being public wealth, should be publicly available.
Data that were sourced from the public, should be public. We need this legislated, or wealth inequality and fiefdoms will keep growing with the historically known outcomes.
Didn't Google take it down because of copyright? With the rise of paywalls it became a backdoor around them, but Google doesn't actually have a license to republish the copyrighted text they scraped.
It's not a backdoor is the sites intentionally served Google a version without the paywall so the indexing got "better" than what real users got. The sites kind of dug their own hole here, Google wouldn't even sit on the content unless it was offered up to them in that way.
Unfortunately copyright law allows them to give something to one person without giving it to others _and_ to sue to the former person to block them from giving it to the others.
Just tried Yandex again, after a long time, and indeed it was a blast from the past, in a positive sense! So many real search results, which transports the feel of an authentic mirror of the web contents. Default search engine from now on! Thanks for the hint.
Yes, it’s been impossible to trust Google since the “Jigsaw” project, and especially when they started blanking out search pages on trending topics with messages like “these results are changing quickly”. It’s clear that they have moved from being a neutral arbiter to pushing a perspective.
I refuse to use a search engine that sees me as cattle to be steered towards “preferred” information sources. I specifically enjoy reading fringe material and will continue to seek and find it, thank you very much.
> There are still illegal streaming sports and movie sites everywhere (who knew) and all other seedy corners of the internet that have been neatly erased by Google
This is such a strange position to take. Apple doesnt allow all sorts of apps on the appstore, but that is never said as "Apple is erasing illegal streaming". Google is simply not showing the results. They are not taking down, or banning, or doing anything to the websites.
The pain being made is: If you use a search engine as your eyes to see what exists “on the internet”, then, absolutely, whatever Google hides from its results or fails to index is “erased” from “your” experience of the internet.
I see where you are coming from but the argument would be stronger is Google Chrome refused to open illegal streaming sites etc. AFAIK, there's no such restriction.
(I dont agree to this but ...) using your argument of " If you use a X as your eyes to see what exists “on the internet” .. " - we should all be mad at Apple. I use iPhone as the primary device to access apps, and they not just hides but actively ban and cut whole swathes of developers.
If I ask Siri to give me a link to illegal streaming site and if it refuses, is that cause for concern? I'd say no. Infact, I dont expect it to give me that and I get it. Same for Google Search in my humble opinion.
I see piracy sites just fine on Google. Almost always the top result. That's why I go there to find the next domain after a previously working one gets shut down.
If you use an iphone then yes, apple is indeed "erasing" illegal streaming in the same sense that google is. We can be nostalgic for the old internet while also recognizing that at least google is well within their rights not to serve up results that fall outside the bounds of the law. (Apple not so much. When you gatekeep the hardware platform I think you're ethically obligated to act as a common carrier. Unfortunately the law doesn't require that.)
> Apple doesnt allow all sorts of apps on the appstore, but that is never said as "Apple is erasing illegal streaming"
Because they are not "erasing" anything, they are refusing to provide a platform that enables the direct (app whose intended purpose is) delivery of illegal content to you.
The difference being one will not facilitate the activity, while the other, is actively suppressing it. And thats what is meant by erased.
Your question in other comment If I ask Siri to give me a link to illegal streaming site and if it refuses, then that would be the better comparison. Then we could say "other seedy corners of the internet that have been neatly erased by [Siri]"
If the companies weren't trying to destroy trust and take everything away from everyone these things would never have much traction but now they deserve to have more than ever.
People have this fixation on privacy, when there's so much more going on, potentially way mor important.
Google's strong position in AI gives them bragging right to attract more companies and get them to actually pay, subsidizing your use. It also lowers competitors position as you're not touching them while we're on Gemini. It also fortifies their position in the future ad market.
Being second or third in AI usage is worth a lot, one's private data matters very little in comparison.
I won't be first, second or third in the AI race. I'm not even participating in that race. I don't care who wins. However my private data are mine and I do care about them.
It's Google and those other companies that may not care about my privacy, because the AI race is more important to them.
If AI is going to cost too much, I will use a cheaper one (Ferrari or FIAT?) or if every AI will cost too much, no AI all and there will be billions of people like me.
Everything uses energy. AI is uniquely bad due to its scale. Comparing AI inference and training with posting a one-line comment on a website is like comparing wildfires with candles because they both produce heat. It's true but useless as an argument.
Plus soon we'll be paying based on secret parasocial influencing agendas.
It's now possible to put a complex spin/bias on whatever the system shows (or whatever it chooses to bury) in a really easy and scalable way.
Ex: Dairy Association pays, and suddenly results about bones are just a bit more likely to show something about the importance of calcium that many people get from milk. Fraternal Order of Police gets involved, and now "shot by police" everywhere gradually morphs into "was in an officer-involved shooting".
... And of course "Google making it harder to see URLs" becomes "Google taking bold steps against evil scrapers."
Me too. I don't connect my phone to my Google account. I understand that the Play store is enticing, but there are other options. Some of which are actually better.
This is interesting. Do you consider Yandex absolutely safe to use (wrt their possible connections with the Russian government, and current war in Ukraine)?
Asking this as a geeky end-user from Estonia -- I wonder if I would be considered "suspicios" by local internet service providers, our govt/police etc if I did a lot of searching via Yandex instead of e.g. Google these days.
(For us in Estonia, I suppose "everything with possible Russian ties" seems suspicious these days, unfortunately with a reason. As a side effect, this also creates a hesitation to look into interesting projects like ReactOS, the old-dos.ru repository, etc. Russia, obviously, inhabits lots of extremely talented coders who still have the skills and mindset to push older hardware to its limits, actually still write stuff in assembly for that, use the reviving DOS distros, etc -- lots of very interesting, creative coding done there, I guess. But, I hesitate visiting their webpages due to Putin's aggression in Ukraine, them being pro-Putin, and possible spying related to all this. It's a sad state of things actually. Sort of like a loss of computer-cultural ties.)
As a general rule, if another country has an open war with your neighbor and openly says that you may be next, don’t put in private data into the sites controlled by that country.
Ok, maybe not the war *with a neighbor* (yet), but US has an open war with Iran (and has bombed a few other countries recently also) and it also threathens to annex the neigbor to the north plus another island country.
Sure, americans might consider themselves to be "the good guys", but the rest of the world doesn't.
From where in the rest of the world are you? Did even you talk to anyone living in countries neighboring Russia?
Sure not everyone benefitted from US but a ton of people/countries did and still do.
US never felt a need to build walls to prevent their citizens/allies from leaving. Russia and the others very much so.
With US the track record may be mixed, but a ton of countries from Europe and Asia benefitted tremendously. Russia otoh doesn’t have a concept of win-win, they tend to exploit even their closest allies, which anyone living in Baltics/East-Central Europe can tell you.
What bombed country benefits from US? Afghanistan? Iraq? Iran? Yemen? Venezuela? Syria? Yemen? How do they benefit? Destroyed infrastructure is somehow helping them? Americans stealing their resources if they decide to occupy the country? How? Win-win in iran how exactly? US just proved to the rest of the world that you need nukes to be safe from americans, and even then they'll sanction you to hell if possible, just because you're not bending over for them.
I'm from the border of central europe and balkans, I don't benefit at all, it just costs me money to pay for a van full of soldiers every time americans decide to occupy some country half the planet away.
Mmmh this change doesn't seem about gaming the index but rather about scraping the index. Which affects in a negative way only Google. And actually would affect users in a positive way because it lets other companies create their indexes more easily (at the expense of a poor mega-corp, indeed)
So how is that AI slop garbage useful? It often lies, aka "hallucinates". And I operate within this Google-controlled auto-spammed text of lies when I use its horrible AI. The last thing I want to do is make Google even more powerful; at the least now with Google search being so crap, alternatives would be well desired, but oddly enough I also get crap results using e. g. DDG, so I think in part the whole www stack must have become more crap too, quality-wise.
I actually don't see it as there are extensions that eliminate this slop spam, but
I never found it useful for anything. Other than waste my time, back when I still saw it. Thankfully I no longer see that AI slop spam due to these extensions.
> Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters".
As long as you don't search anything related to Russia itself, or to Russian interests elsewhere (like their invasion of Ukraine). Then it's very heavily biased, priority is given to state ran propaganda mills, independent media is hidden from results, etc.
And before anyone starts thinking about whatabouting, Yandex are based in a country where journalists are openly and publicly assasinated to intimadate. It would be delusional to expect any sort of press and related (like search engine or aggregation) freedom, or to try to compare this to anything in any other developed country.
A Soviet citizen and an American are sitting next to each other on a commercial flight.
The American turns to the Russian and says, "I have to hand it to you — your state propaganda is very impressive. You really know how to shape what people think."
The Soviet smiles, nods, and replies, "Thank you, but it's really nothing compared to American propaganda."
The American looks shocked and says, "What are you talking about? We don't have any propaganda in America."
Nobody is denying the existence of American propaganda, but it's an absurd false equivalency to pretend it's comparable to Russia - the country where there are multiple Wikipedia articles with lists of murdered journalists and there is a day of remembrance for murdered journalists. The country which sends assassins to kill dissidents abroad. Especially when we have the Mitrokhin archives since the 1990s in which a KGB archivist painstakingly tried to warn us about the firehose of shit method. Do you know the conspiracy theories about AIDS were started by the KGB? This is the kind of shit we're talking about, not Americans jacking themselves off on their anthem thanking troops for their service.
So, I decided to check, and the first links that show up when I search for "war ukraine" in yandex.com are Google news are from rbc.ua and bbc.com. When I search for it in yandex.com, the first link is CNN, the second google news, the third is indeed TASS, but the fourth is novayagazeta.
I then picked "bucha" as the query most likely the be affected by propaganda and censorship. In yandex.ru, there are propaganda outfits in the result (something called ruwiki.ru comes third), most links are of fairly direct accounts of the massacre by independent media. The AI summary says the town was "occupied" by the Russian army and "liberated" by the Ukrainian one and "После отступления российских войск в городе были обнаружены многочисленные свидетельства массовых убийств мирных жителей." On yandex.com on the other hand, the first page of results does look fairly propagandistic, including "globalresearch.ca" and "donbass-insider.com", which seem like propaganda outfits. But it also does include accounts from novayagazeta, al Jazeera a video of killings from Radio Free Europe.
While they do put either own propaganda, neither Western, nor Ukraininan nor independent media is not hidden from the results and it's not visibly de-prioritized.
Of course, I wouldn't discount the possibility that the results would look quite different from Russian territory. And, if anything, half the reason why Russian propaganda is so effective is that its self-aware and capable of subtlety when needed. "Fuck you if you can't handle the truth, this version of Biden is the best version ever." would never happen there. Unlike Scarborough, Soloviev knows exactly what he is.
Open and constant suppression of information is not the regime's usual strategy in information space (though, they'd ratchet up the level of control when they deem it necessary), which is why I knew the claims here wouldn't stand up to scrutiny.
I explicitly say, that I won't be surprised if the results change from Russian territory. I already put way too much effort to check the validity of an internet comment, but if anyone is interested, they can check.
there was a source code leak and you can see some examples there which make the president's image more alluring. e.g. if you search for "Putin" they add an invisible "-crab" to your query
that's the most innocent example from the top of my head. There probably are many more serious censorship issues
Yep, plenty, and tbf it's entirely logical. The people working at Yandex do not want find themselves dying of a nerve agent or polonium, which is a real and acute danger for people who displease the ruling regime.
>It would be delusional to expect any sort of press and related (like search engine or aggregation) freedom, or to try to compare this to anything in any other developed country.
Israel is responsible for killing more journalists than any other country. Google has profited by providing services to Israel in the ongoing genocide.
I was waiting to see how long it took for someone to "make this about Palestine" - their strategy is to inundate all discussion with the topic.
Over half the journalists in the Gaza Strip are affiliated with Hamas. Many of them infiltrated Israel in 2023, and some of them held Israeli hostages in their homes, including children hostages.
Hamas claims that the journalists are Hamas, not Israel. And I'm not declaring any journalist "good" or "bad". I'm stating that a Hamas member is a valid target. Him moonlighting as a journalist does not change that.
Every claim need to be proven and everyone deserves a process. In a democracy of course, it may be different in a fascist state committing a holocaust.
That site seems completely unbiased and not at all fed by the propaganda coming out of Israel. Thanks for finding such a reliable and trustworthy source.
Parent comment is implying that an ally (Israel) is performing a holocaust. In fact, there have been a good dozen genocides in the past century which I would agree fit the term, but not the conflict that Israel is involved with.
The anti-Israel and anti-US side typically take an argument or incident from Israeli or US history or culture and invert it. A good example is the common Hebrew expression "No other land" - we have culture and songs about that, as well as jokes and it is a common expression in everyday speech. After October 7th, as part of the propaganda campaigns against Israel, that phrase was used as the title of an anti-Israel movie highlighting that another population has "no other land". Inversion of our culture.
Holocaust inversion - accusing Israel of perpetrating a holocaust - is this particular example.
To other readers: Parent is suggesting that the 6000 Gazans who breached the Israeli border on October 7th 2023 and killed 1200 people, mutilated bodies, beheaded humans and dogs alike, raped women, and took over 250 hostages including infants... Parent is suggesting that these people were undercover Israeli agents performing a false flag operation.
> And before anyone starts thinking about whatabouting, Yandex are based in a country where journalists are openly and publicly assasinated to intimadate
Thankfully Gary Webb is here with us today to laugh at this
I turn it down straight away for being Russian and I'm suspicious of anyone that mentions it positively.
Russia is in a cold war with democracies in general and a hot war with Ukraine. It has an autocratic dictator as leader and state control of media.
It has a well documented propaganda campaign with the goal of destabilising other countries through agitation. We even know the address of the building in Saint Petersburg it's run from.
Why would anyone try to assess Yandex's results? They could be like a slot machine serving good results 78% of the time and slipping propaganda in when least expected.
Why bother with that when there are good, free alternatives?
Couldn't we say the exact same thing about any US service in this day and age?
Trade war with half the planet, hot war with Iran. Autocratic dictator as leader and complicit media. Well documented propaganda campaign and history of destabilising countries around the world. Known to compel tech companies to spy on foreign citizens...
I have, and it has never been good for my purposes. In English, I just can't find things I need, in Russian, I find things I wish to never see again, tbh.
What is your experience with it?
Edit: The best search for my purposes historically has been on Pinterest. Unfortunately, not anymore. I normally search for visual artifacts (photos, art, schematics, etc), out of all search engines Google is still the only one that is somewhat functional for me in this regard.
I tried DDG, Kagi, Yandex. None of them worked for me.
The quintessential HN comment questioning the obvious.
Yandex does Putin's bidding much much further than any western search engine ever meddled with their results for political purposes. Yandex employees' lives would be actively at risk if they didn't.
Sometimes I really wish a hard reality check would meet the people making such comments from the comfort of their programming chair or from thousands of kilometers away.
> Yandex does Putin's bidding much much further than any western search engine ever meddled with their results for political purposes. Yandex employees' lives would be actively at risk if they didn't
Does media manipulation score higher/lower on your moral index than the widespread domestic spying the US has forced tech companies to be complicit in?
Not that media manipulation by the US government is exactly an unheard of event (at least as far back as the whole WMDs in Iraq scandal).
> Sometimes I really wish a hard reality check would meet the people making such comments from the comfort of their programming chair or from thousands of kilometers away
Look, I'm American myself, and have as keen a dislike of Russia's foreign and domestic policies as the next guy. But it is at best wildly hypocritical to suggest that US tech firms are not subject to government interference, in a manner that may be detrimental to their customers.
> But it is at best wildly hypocritical to suggest that US tech firms are not subject to government interference
It is very different degree of subjection where we talk about country that have laws, politics, сourts, journalism, all that stuff, and country, where if you dare even to complain about some wild highly illegal stuff that the kgb demands corporation that you working for to do, you end up getting 30 years in prison for treason, falling out of the window, getting poisoned by some chemical weapon, and even then you will be considering yourself a lucky man because nothing very, very bad had happened to your family.
I did an interview with Google around 20 years ago, where they posed a challenge involving tracking which specific search results people click. It's obvious in hindsight the solution required rewriting all the urls to redirect through their servers. Note this was in the days before they already did so as a matter of course.
I failed to gain traction on the problem, because to me the very idea of doing such a thing was too reprehensible to seriously consider. It broke an unwritten contract between the company and the user's expectation of how websites worked. You expect to be able to do things like right-click a link and copy the authentic URL, or hover to see where it wants to take you. The notion of obfuscating the link beyond easy recognition and polluting it with tracking markers felt misleading and, well, evil. A move that would mainly only benefit Google, and not it's users. I (quite mistakenly) presumed this opinion would be obvious and self-evident to anyone who spent enough time around the early web to understand its norms.
I explored other ways of achieving the goal, but it clearly wasn't the answer the interviewer sought.
I'm more seasoned now, and experienced enough to say with confidence the approach was wrong. This may seem like a small thing, but a series of misteps and chronic failure to adequately advocate for users is what has led us to the toxic waste dump that so much of the Internet has become today.
I'm really glad to have fresh alternatives (like Kagi), and can't wait for the cultural zeitgeist among developers to swing back around to valuing users as human beings and living up to the trust they place in us.
I had some fun in my grad days building a private information retrieval (PIR) scheme so the shim server can do its job without knowing which link you clicked: https://github.com/pncnmnp/shimmey
I switched to DuckDuckGo around 2018 or so, which was when I realized that Google's results quality has deteriorated so much that I won't lose anything of value by doing so. This thread is how I learn about the atrocities that Google Search is committing these days.
Which is an elaborate way to say: you don't have to live like this.
I don't. At this point I just ask the oracle in natural language. There's so many to choose from. Why spend hours wading through only vaguely related things? You can ask it for a direct answer or you can ask it to teach you about the topic.
Sometimes there is joy in discovering things for yourself. Also, the oracle currently relies on humans publishing new information, what will you do when the incentive is gone and the oracle has nothing to harvest?
> Sometimes there is joy in discovering things for yourself.
Agreed. And webrings still exist in certain corners of the internet.
> the oracle currently relies on humans publishing new information
New? When I ask it how to do something with a piece of software it's perfectly capable of answering me by means of direct interrogation of the source code.
I'll grant you that it owes ~all of its knowledge of the world at large to having been bootstrapped using the more or less complete body of humanity's published works. However I think that was merely the quickest path to bootstrap it as opposed to the only one.
> This may seem like a small thing, but a series of misteps and chronic failure to adequately advocate for users is what has led us to the toxic waste dump that so much of the Internet has become today.
I think it is both monopolization and cheap money. Something in economy favors big companies and disfavors competition.
Cheap money means that a selected company can run at loss, kill competition and then enshittify to start earning. By that time it is too late and too easy to buy or destroy smaller competing company. And you get access to funding by being charming to VC, by being the kind of sociopath they like to see.
Capitalism works when there is a competition. Not when there is an oligarchy.
I was suspicious when they started obfuscating URLs in their own browser, then on their SERPs, and now this...
For many years, I had my filtering proxy rewrite the URLs in the way mentioned in the article.
Almost exactly a year ago, Google stopped working without JS. I stopped using Google.
Now they're upping the game, and as the article (which is a bit of marketing itself) admits, those who have the resources can still blast through these obstacles while those who don't are locked out.
Since the article brings up "AI scrapers", I'll just point it out as being the latest scare-tactic for coercing people to give up the privacy, anonymity, and (browser) freedom of an open interoperable Internet.
> Sometimes, these redirect URLs take a perceivable amount of time to load, which is very irritating.
Great. On top of my on-going battle with Windows + Firefox + DNS/TLS resolution sometimes stalling for seconds at a time, another few second server-side stall is introduced.
I swear that every day modern computing scenarios get slower and slower instead of snappier and snappier.
Wow - this exact same bug has been happening to me too. I gave up on troubleshooting it after the first few attempts came up with nothing, assumed it was just unique to me.
I searched for this a few months ago and found some mentions of this bug, but yes it affects me too. Can stall up to 20 seconds+ sometimes. Chrome is fine.
> Sometimes, these redirect URLs take a perceivable amount of time to load, which is very irritating
I do a bunch of work in remote areas with low-ping/low-bandwidth networks. The “hold on a skimmer while we round trip to Google” dark pattern makes the service unusable for me when I’m on-location.
The biggest problem with this is if the target URL doesn't load but also doesn't quickly error out, like e.g. many .gov sites in Europe (seems like they are just dropping traffic from non-US IPs).
Now you can't load the page and can't easily (using only the browser UI) get a link to paste into archive.org or archive.is to read the page.
Try it in a private browsing window - if you are logged in, Google will link directly to the result, but if you aren't logged in, it appears to be redirecting through their /goto endpoint
The link you followed when you clicked hasn't a direct link for years, decade afaik (they mangle so they can see what's followed). The page used to show the direct on the search text but now it shows some stand in for it - sometimes. You can see the direct link on the bottom of the screen when you hover - sometimes (and sometimes you see a mangled link). Sometimes the google link contains the original link in the center also[1].
The situation seems to vary from result to result even on the same page of the same search - at least on the test search I just did. You can figure out what happening to an extent but this very inconsistency seems to speak to a dystopian quality to today's information gatekeepers.
I've just checked this again using a google account where search result pages are still following the old behavior: It looks like the 'href' attribute is the direct link, and the 'ping' attribute is the /url redirect link you are referring to. So it looks like it is actually sending me to the direct link, it just also requests /url at the same time in order to log the click. This means the user was not waiting for the logging/redirect request to come back.
Which is exactly the game theory that was predicted when some browsers started ignoring <a ping> to "protect privacy". If your browser supports ping you get ping, otherwise the website gets the data anyway but with a worse user experience.
While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be to build a set of sites of most interest to you and then prioritize those in crawling.
While typing this out I remembered https://en.wikipedia.org/wiki/Google_Search_Appliance which I never personally used but shows feasibility for the idea. I'm pretty sure one of the newly-announced Macbooks is more than up to the task of matching GSA's offering.
Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.
It would be interesting to see a decentralised, residential collective that builds and publishes an index. There are surely enough interested people on HN alone that would be willing to run software at home to scrape a small slice of the internet.
The majority of websites try to do that but they do not catch as much traffic as they think they do. A starting point for a scraper is to run it on your home connection in an undetectable web driver framework such as zendriver.
This was a knee jerk response to the first paragraph. They weren't talking about a general crawler, but a subset of Wikipedia, stack overflow, programming docs and github. You can download archives of all of those except github, and github could be queried using the api or GH cli
For some of that, you don't need to crawl. Wikipedia offers database dumps which you can download in one go. Lots of programming docs are managed in repos, so you can clone the repo instead. Even stackoverflow seems to have a snapshot dump (https://archive.org/details/stackexchange).
SearXNG configured as in the OpenWebUI docs is pretty cool. My "Hello World" with a new agent framework is teaching it to use SearXNG. Hook this in as a tool and the agent can answer a lot of questions.
SearXNG is more of a metasearch, the dude who wrote it pops in on here and is working on a cool sounding project that is more like a local personal search engine, I forget the name, but I've been meaning to check it out.
SearXNG has been working pretty well for me. I had an agent write the MCP then do a couple passes comparing to server side LLM web tools and exa and tweaking and it works pretty well. I also added scrapling for fetch which covers pretty much everything but sometimes is a bit context heavy.
Wayback Machine full archive is like less than 50PB. Let's say you could strip multimedia and remove every patterned data to compress that into about a petabyte. The per-bit cheapest disk right now is consumer Seagate 24TB(SI; 21.8TiB usable) at ~$500, or around $1200k for just the disks.
Doable if you had couple million dollars to burn. Cheaper than private jets new.
Seems like the most practical path is some kind of distributed peer to peer contraption where you can allocate some storage and optionally participate in crawling.
I imagine if set some constraints you could get index size down quite a bit but I still suspect it'd be hard keeping up with content churn.
Edit: Looks like enwiki bz2 is coming in around 46Gi which isn't too bad considering the amount of content it contains.
It's still kind of an open problem. There are partial solutions, but not yet really an integrated one. There is Hister: https://hister.org/ that builds some sort of drive-by index of what you're browsing anyway. And there are "true p2p" solutions like YaCy: https://yacy.net/ but it takes ages to crawl the open web (and lots of storage). There's still a lot of optimization to do in this space.
Plausible. Figure 100 GB each of search index for Stack Overflow, Wikipedia, and GitHub issues, then add a dozen more for docs of all your favorite techs. So maybe half a terabyte. Download and build updated dumps of those once every week or two, and it'd work pretty well. Impractical, but possible.
That's what a redirect is, reading the value of the location header and then requesting it. How do you not follow the redirect by reading the location header? Once you've made the request to the /goto url, with GET or HEAD, to get the location header, google knows you're interested in whatever it is putting in the location header and can assume you're going to go there, if you're letting the User Agent (curl or the browser) go there for you or not.
When Google stopped paid API search a few months ago, I looked for an alternative for my agents that I felt would be sustainable (one-time setup, then out of my mind). I quickly excluded SERP as I feard Google would pull exactly this type of shenanigans to cut them off.
I somehow found Mojeek and settled on it. I had never heard of them. Unlike Kagi, their business model is ads (so they hold no particular moral high ground). But they have a cheap, working paid API.
What I was astonished by is the quality of the results. For my uses, it's undistinguishable from Google. The conventional wisdom is that web search is a Google-sized problem. How did those obscure Brits pull it off?
I’m guessing that the knowledge of the techniques to support large-scale web search have diffused out of Google - it has been a few decades after all. Not to mention that the distributed system knowledge that used to live only in Google was either published by Google or cloned in other projects, usually made by ex-Googlers. And you can now rent capacity at scales that 20 years ago required Google to build lots of their own data centers.
Also, there is now a use case for paid search APIs - LLMs and agents - that essentially didn’t exist a few years ago. Not sure why Google hasn’t leaned more into this, but probably a combination of Gemini and Ads interests have combined to view their search index as an increasingly valuable asset, when if anything it might be corrupted already by these interests and therefore be less valuable.
I've been using DuckDuckGo for years now, ever since it became noticeable that two different people searching for the same search term would get two different results back from Google. Meaning they were no longer completely reliable: they might show one person a result that they hide from the other person by burying it on page 3 where few people ever look.
DDG's search results have been poorer recently than they used to — I often see completely unrelated results (to the point of my saying "Why in the world did that come back as a search result??!?") starting from page 2. And yet, I still use them, simply because they aren't Google.
That somehow hurts me to hear more than anything. I remember days spent in my youth trawling through Google search for anime fandoms and homework answers. Joke used to be that beyond page 1 is the definition of desperation and in the 00s that was definitely true but I discovered a lot of cool distractions in that desperation. This feels like they killed Google Reader again.
Don't know how modern this is, but I remember being disappointed that despite Google claiming it had billions of results, you could only see a few pages' worth.
Looking now, I've just noticed they no longer give a count of results.
I don't know how global consistency is actually useful, but I can easily think of ways it is less convenient.
For example, DST starts and ends on different dates in the UK and the US. Which dates should google return when someone from either country searches for just "DST end date?" Someone lives in Orange County and searches for "Orange County Sheriffs," which one of the eight Orange Counties should google return?
These are both examples where a localized — not even customized, just localized based on IP addresses — search results will easily help reduce headaches.
I wouldn't mind that one. That a man in the US and a woman in Japan would get different results for a search for "sushi restaurant" is perfectly reasonable (even if the woman in Japan was searching in English).
It's when two people in the same neighborhood got different results for the same search that I said "wait a minute, they're personalizing search results now for ad-targeting purposes" and ditched them. The potential for them to deliberately hide things from you was too great.
When I search Python I should get the programming language but when my neighbour searches Python he should get snakes. That is a valid use of personalised search.
Good example, which is why both you and the other commenter gave it. And if that's as far as it went it wouldn't have bothered me so much; it was when conservative people I know were getting different results from liberal people for the same search terms that I started to realize the potential for an echo chamber, and decided it was a bad idea.
It's good for your mental health to be exposed to arguments you disagree with. Many times they will be bad arguments and you can dismiss them, but it's good to know what the people who disagree with you actually believe. And sometimes, just sometimes, you might see something that makes you realize that it's your own worldview that was mistaken, and adjust your worldview to better fit reality. If your search results only show you echo-chamber results that already agree with yours, you'll rarely learn the areas where you were mistaken. (And nearly everyone is mistaken about some things; the person who is right about everything is a rare creature indeed, and you should never gamble that you're that person).
The age of internet search is over. The age of Cloudflare has begun. It wouldn't be possible to build a search engine now if they wanted to... and there wouldn't be anything to search for anyway. The non-corporate internet withered into dust and blew away in the wind.
If you could find what you want, how would they ever sell you what they want you to buy? And I'm not just talking merchandise, though that too. Your political narratives, your values, opinions, everything. And everyone likes it so much they just sit there scrolling and swiping and tapping.
I'm hoping for a future where we create static HTML pages again styled with a bit of handmade CSS, because we're so tired of bot attacks and long loading times. Then suddenly a Cloudflare network becomes absolet.
people have been putting their fuckass blogs with 0.5 visitors a day behind Cloudflare long before the increase in bot traffic. and that increase makes fuck all any difference for them anyway.
It's probably cheaper to switch your ISP to one that your target site does not block, than to pay someone else to buy that connection and proxy all your traffic through them.
I think curated directories from prehistory might make a comeback eventually, maybe combined with a search engine that indexes only the whitelisted websites. ironically, we can use LLMs to filter out AI slop by measuring the signal-to-noise ratio, which is atrocious for AI slop and ESL slop that preceded it.
It means they're capable of burying news stories that would contradict your worldview and pushing news stories that support your pre-existing biases, leading to more engagement from you (a win from their point of view) but also burying you in an echo chamber. And unless you were in the habit of doing the occasional search in Incognito Mode, you wouldn't know. (And even then, they probably would be able to put together enough clues to figure out your identity even without your Google login cookie).
Yes, the fact that they could do it does not prove that they were doing it, not right away. I ditched them as soon as I found out that they could do it, because I was absolutely certain that eventually, they would end up doing it. And I wanted neutral search results, not biased ones, even ones biased towards my own point of view.
I subscribe to The New York Times, if I am searching for a news event, I’d appreciate a site that’s not paywalled and I trust be the first result if it is reasonable.
How u reliable personalised search results are depends on who controls the personalization and for what purpose.
I would love a search that lets me ban results from certain sites. I don't like search that's showing skewed results for reasons I don't know or control.
"30. Copyright holders have authorized Google to implement access controls like SearchGuard for the content they license to Google, and in some cases insisted that Google do so. Googles authorization takes many forms. For example, Google has an agreement with a prominent licensing partner that holds copyrights to millions of works that it licenses Google to use in its Search results. Under the parties agreement, versions of which date back to 2017, Google is not only authorized, it is obligated to use commercially reasonable efforts to safeguard the licensed content against unauthorized third-party access. Other license agreements contain similar obligations. For example, another major content provider requires that Google ensure the content it licenses will not be available for download by third parties, thereby authorizing the implementation of technical access controls."
"31. In other cases, Googles authorization to implement access control measures like SearchGuard is part and parcel of the grant of licenses themselves, as Google and its licensors recognize that the value of the licensed rights would be undermined if others were free to access, take and resell the licensed content without restriction. For example, Google has a licensing agreement with Reddit, under which Reddit licenses Google to use the copyrighted content of both Reddit and its users in Search Services."
"32. Googles licensing partners have also expressly requested that Google prevent unauthorized access to licensed content. For example, when Reddit suspected that scrapers like SerpApi were accessing, taking, and reselling the content that Reddit had licensed to Google, it specifically asked Google to employ technical measures to prevent such unauthorized appropriation."
But this does not account for material that is not covered by the "license with a prominent licensing partner", its license with "another major content provider" or its agreement with Reddit
Google not only uses SearchGuard on SERPs containing links to the content covered by these licenses, it uses SearchGuard on _all_ SERPs
Google needs more than a "goto" update. It needs to update its terms to require _all_ copyright holders for the materials it has indexed and cached to give Google authorisation to use "technological protection measures" to deny access to certain members of the public, e.g., Google's perceived competitors including any Google user who "searches too fast"
They are allowed to scrape everybody else but get their feelings hurt when they get scraped....oh yea they respect robots.txt. Guess what, scraping everything that is public is legal.
How do you even measure GEO, sometimes I try searching for my app and sometimes claude recommends my competitor that has 10x less reviews and is objectively worse saying is the most popular.
Because Google search isn't a hypertext document, it's a web app that presents dynamic ephemeral content to the user. There's no need for stable hyperlinks, since either you will click them within a few seconds, or they will disappear forever, and it shouldn't be a problem to indirect via Google when clicking them, since you just made a request to Google to get the link in the first place.
Hmm that's actually pretty clever. They can serve each result page slightly different encrypted links and it should be obvious right away if it's a SERP bot (trying to grab a page of links) or a human that just picks a few here or there.
I wonder if this would also work on other sites getting hammered with bots. Allow each anonymous user 1 "real" page load then turn the rest into encrypted links that the web server can decrypt. If a session cookie with reputation exists, stop screwing with the links.
Kind of annoying but it'd allow tracking if the same agent/bot is churning through IPs/User Agents.
It's rather frustrating that after they crawled my website to use for their AI without my consent, and using that AI to slowly strangle my traffic, they're not happy that _other people_ are doing the same to them.
>”This is what capitalism's all about. Take as much as possible and give as little as possible if you want to win. Ideally, give negative amounts.”
Your description is vague enough to encompass every system that I’m aware of, from feudalism, through to mercantilism, communism, and capitalism. It may best fit communism and feudalism, both of which are most notable for fostering negative-value-added firms.
It breaks the social contract of the web. My user agent should be able to tell me where a hyperlink goes without first clicking on it, and perhaps act differently based on that information.
Now, all Google search results appear to go back to Google.
TBF this is more a browser problem than a website problem to my mind. Redirects shouldn't be handled in such a cavalier manner. I should be able to configure the browser to stop and confirm the destination any time the domain changes without direct user interaction.
2. It is non-bypassable tracking of every single click.
3. It does not allow you to inspect the URL to see where it goes to before visiting.
4. It breaks the feedback signal of extensions which redirect sites. E.g. Fandom wikis are shit so I have them redirected to equivalent much better wikis like. But now any such redirection has to be done post-click tracking meaning Google still believes I want to see the Fandom site.
5. It breaks extensions which hide certain shit search results based on URL.
That's just Google not caring about your ISP because you live in the global south. Everyone gets this. US citizens get the exact same problem but Google whitelists their networks.
Aside from the list of current concerns that others have covered well...Google is very good at "boiling frogs". That is, rolling out unpopular things in phases.
This could be, for example, step one in the return of AMP, but with a new twist. Where google conveniently returns the content of your website, without directly sending the user to your website. With whatever changes it chooses to make.
Personally, I put a lot of effort into writing my blog posts, some of which have been really well viewed by humans. While mine is effectively a "hobbyist" site, it's nice when readers look at other blog posts as a result of reading the original one. Google's approach ensures that readers get a very narrow view of my content, so everybody expect Google loses out.
Aside from the privacy concerns and inconveniences this brings to Google users, it matters for users (of services) that rely on SerpApi too. Kagi for example has started showing these goto URLs in search results.
Well, that's rather the point, right? Kagi and other micro search engines sell a product that just resells Google search results while telling people it's a premium product better than Google. Obviously Google isn't pleased about that.
Yeah. I might be wrong but I think they only served direct links for a relatively short time in their history, early on in the 90's and in recent years with ping, which they used to track clicks anyway. At least half of their history they used either the 302 redirects or onmousedown link rewriting (which was terrible). And I'm not even starting on AMP.
In this case they assume correctly that it is enshitification. A search engine, that doesn't understand what hypertext and the web are supposed to be is shit. There are no good reasons to not have proper links. Whoever creates such websites is having other motives or doesn't know good web development. In the case of Google employees I have to assume the former. They are not acting in the best interest of the users, and therefore enshitify Google search.
Well, I hope ultimately this will lead to even more user loss for them. Fortunately, I myself don't have to suffer due to it, because I degooglified my life.
It's sad that instead of searching things other people put up, we're basically asking sam or dario oracle to tell us the truth. The people should be furious. But we've internalized this idea that they are somehow better.
Not better, just much _much_ more convenient. And it doesn't have to be a frontier lab. I'm happy to ask any oracle that returns sufficiently good results which is an ever increasing number of them.
If you were scraping only Google with all the IPs you can get, then this change really slows you down.
If you're trying to fight scrapers on a small site, delay links can only flatten bursts. If bots can only scrape at human speed per IP, they can just scrape 100x as many sites at the same time. Once every bot operator does that, total traffic will return to the original level.
It's primarily relevant because it makes scraping search results much more expensive, solidifying Google's effective monopoly on Internet search.
Google has previously tried to prevent scraping of search results using legal means, but courts correctly think that scraping of Google's search results should be legal, just as Google's scraping of the whole Internet is legal. This is Google's reaction to that.
This is the best explanation. They’ve been doing the same in Google News. Each entry comes not with a URL to the source, but with a hash. To resolve it, you must send requests to Google’s servers. Anyone who wants to create a list of URLs of sources automatically can therefore be blocked by Google now on two levels rather than one - the search for a list of results, and identifying the source URL for each result.
In effect, they’re removing attribution from the content they quote from other people’s websites. It would be interesting to see if courts object to that. It is one thing to crawl other people’s websites and display snippets of their work as your search results when each result is properly and transparently attributed. But if the text is quoted and the source is not there alongside it in plaintext, replaced only by a vague promise that, if you ask, we may or may not tell you where this piece of content is from, that is a very different deal.
You can opt out from Google scraping you though? In theory you can opt out of anyone scraping you (if people were well behaved). Google should get to opt out of being scraped too.
> You can opt out from Google scraping you though?
You can't. Google will ignore robots.txt in some cases (e.g. "The REP isn't applicable to Google's crawlers that are controlled by users (for example, feed subscriptions), or crawlers that are used to increase user safety (for example, malware analysis)").
robots.txt is just a suggestion that Google loosely follows.
Well just as a regular user, I think it is pretty annoying because if I look up anything on Google while in Incognito Mode and hover over a search result, I can see that maybe the top result is maybe Wikipedia, or Instagram, or some other less-known website depending on what I'm searching for. Now, that's all very obfuscated because I don't actually know where I'm going to land for sure.
I copy a tracking link from google in my reply here. You click on the link. Google now connects me and my search session to you and knows where I gave you this link from. If I get the original url, all google knows is that I copied that link and nothing else.
They were already tracking everything you click but for example if you want to send a link to someone you can't copy the link from the Google result and send it to them, you'd either send them the Google tracking link or go to the website yourself.
it sounds like it primarily matters if you are a customer of this company, one that is building a search index off of urls scrapably hardcoded (or at least so as to be easily unencodable in non-realtime, it sounds like?) inside google search result redirects. in theory, there could be noticeable consumer user impact, but ... it would have to be a pretty large theory
It's to make it more difficult for scrapers. Google have been able to do all the tracking they've wanted for decades.
Think about it: if you want to set up something like jina.ai you need to build your own index or piggyback on Google. My guess is they use residential proxies to fire requests to Google then scrape the URLs.
Opaque URLs now give Google another gateway to detect circumvention of their anti-bot controls, helping them monetise their search index rather than allowing providers like jina.ai to succeed.
The security implications of this are very severe when you consider the amount of people who google government websites, banking, crypto and others. And google will happily serve you a phishing website either in ads or results.
> the reasons are so they can track who you are and sell your profile advertising.
What? Like they weren’t doing this before? Obviously Google’s telemetry is tracking every link you click regardless; there’s no extra tracking benefit to this.
The reason they’re doing this seems to be to stop competitors from scraping their search results.
It benefits on the other end, if I share the link and you click on it, Google knows you got it from me and from where. People 100% don't click through to share original URLs.
Dont worry in a year or two, Google wont even redirect you to the actual true url, instead everything will be a page with all links rewritten so all http is tunnelled thru them.
Kagi is what I settled on a couple of years ago. I do think sometimes the fawning is over the top, but it is a solid search engine and tended to give me a bit better results than Google out of the box. The real win, though, is that you can give various sites a weight, so the search results will prefer or avoid sites according to your desires. Once I had that going, my search results tended to be much better than Google.
Yep, this. I was delighted they implemented this simple yet powerful idea. I don't even remember when I have needed to go to a second page of search results in Kagi and spammy clone sites which copy stack exchange verbatim are banned into oblivion where they belong. Fuck up once doing that shit, and you are forever out of my search results.
And then there is the very handy assistant, that I can jump into using follow-up questions to the quick AI queries one can choose to use or not to use by appending a "?". Again a simple idea, which empowers the user.
My ongoing concern with Kagi is they always seem to be focused on sidequests like their Orion browser and their LLM-powered Translate tool. Maybe that's interesting for some people, but I can't help but feel like I just want a damn search engine than works.
I also have no interest in their side projects. It doesn't bother me that they have them, though. As you say, as long as the search engine works and the price is acceptable, I'm happy.
Yeah, that's the thing for me: filtering out the SEO crap that Google happily serves up.
Google's actual search results are a waste of time visiting, both for the mindless CEO content and the ad-laden, analytics-happy, javascript-heavy websites.
So I tend to use the AI overview. But plugging myself into the all-seeing corporate oracle, that grew on all the web's content, and now seeks to supplant it seems unseemly.
It's just sad that kagi will likely only be a fringe thing, and google will continue to promote these foul, foul, mindless websites and then supplant them with its AI.
I don’t fawn, I just pay and use-and-forget. It’s almost to the point where it’s a little mental bump when I have to make a browser use it again, setting up something new.
Kagi has a pile of features I’m not getting the benefit of too, I’m sure, because it’s just-search, mostly, to me.
That’s fine; I know what I’m supporting, and I know that I’m the customer and not the product.
Because I wouldn't want Google to know what links I click on. When I used to use Google, and they added such redirects that included the destination URL as a parameter, I'd edit the link to make it just that URL before resolving it. This new scheme would make that impossible.
I use a Firefox extension to rewrite those Google redirects into plain links, to remove Google tracking of which links I click. The extension is broken now
Blocking URL shorteners and google-ad links yes! Personally for me it's also the fact that this is effectively an unresolvable URL shortener, store that link somewhere and it will most likely be dead. Can't copy link anymore and paste it on a notepad or chat app to check it out later as there isn't a guarantee it will load at all (ie: the problem with url shorteners).
Can confirm that when using google not logged in. Now when sharing a link from google search, I won't get the actual link. This will certainly help google's tracking.
Google has gone so bad over the past few years. You only get like 8 results per page. I remember there was a time that I wonder how a site get reached if it ranked on the second page, when I can set the number of results to be 50. The censorship is also really bad, and google doesn't even tell you the results are censored, returning totally nonsense results while other search engines work normally.
I've been supporting Brave search which returns 20 results per page and has other features. It used to be not good a few years ago, but now the results are often better than google's.
This is such an incredibly annoying and deliberate defect. I want an extension that will let me resolve the true url without visiting the site, or even rewrites the entire page to show the url.
I've never heard of this particular SERP provider, but some marketer is very excited they wrote this blog post right now (500+ upvotes on an seo blog post).
And their fix here - they just resolve all the urls- which I suppose could make the service slightly more expensive? But otherwise isn't that what every similar provider/ crawler etc will do and this change will only hurt users?
At least they seem still provide results for my searxng instance. I mean sure, they are horrible but duckduckgo just blocks most queries (and I'm the only person using the ip / seraxng instance)...
Next i'll do is to implement tavilly, exa, tinyfish etc. as search engines for searxng. No agents, no mcp, just their search api endpoint.
I use kagi exclusively because it works. I get the results I need every time, and it's not annoying. I haven't used Google in years. Simply no need. I really hope people keep paying kagi so they stick around.
+1 for Kagi, even if Maps and Shopping and Images aren’t as good as Google (tbh their Images search is good enough 8/10 times) they’ve got the actual web search stuff working quite well.
> Combined with earlier moves like removing &num=100
Removing this made google search horrible to use. I often use command+f to quickly identify relevant search results, but doing it on 10 results at a time is so laborious that I just don't bother using Google search, resulting in less searches and use of other tools instead.
I refer to tools like LLMs, unfortunately (not search engines). If anyone knows of a search engine that returns 100 results, I'm very keen to learn of it
I had tried now but I don't see the goto. Is it possible that has been deployed only for USA users ? (I work and live in UE). If so: probably a vpn can help you for a while. In the future: I think I'll really go for payed search engine.
I... Don't see it? It's the result page right? I just search some random string on Google and the results are all direct URLs. Do they get resolved via javascript after page load and replaced automatically? Or am I looking at something else?
What I see now, and it's been like this for a while, is this:
You get the results and they do have direct URLs. But then, if you do some things with the link, e.g. right click to open it in a new tab, it swaps the URL to the indirect one. The idea is that initially you see a normal link, with a normal URL which will be displayed correctly when you hover the mouse over it, but right before you click it, it's swapped for the indirect one.
So, they have been doing stuff like this for a while and it has been somewhat fluid, because the swapping can occur on different events and I have also seen it load with all the links pre-swapped to the indirect ones, sometimes.
So, yes, what you see may be different and you may get the indirect URLs swapped at different stages.
Becuse they are a monopolist and we can require certain things from monopolists.
Similarly, Bell Labs was kind of required to release transistor for anyone to license - it was a part of social contract that they were allowed to maintain their monopoly in exchange for releasing certain parts of technology.
Alternatively, they could be split up and their indexing division made an independent company selling to anyone on a free market.
It has steep costs, but with monopolists they have other ways of extracting value from the inventions.
Awesome book about the history if Bell Labs - virtually all semiconductor tech we use today was created there (transistors, ics, solar, lasers, fiber optics, telecom satellites…), and they had to license it to be allowed to maintain their monopoly status.
I really like the idea of turning the search index into a public utility. It is one of those natural monopoly coordination problem things. Just quasi nationalise it for economic efficiency. Ofc google can still sell adds against their own ui (like everyone else). Hopefully this move will move that idea closer to reality.
Been using Brave Search now for a while including its Brave AI and aside of sporadic times I never needed Google (albeit Brave Search is slower to Google, you get used to it)
That flag is the result of clicking on the "Web" tab of the results page, so it kinda is a documented feature. All the tabs have their own udm value (e.g. image search is 2, AI mode is 50, etc.).
Whenever I've tried this, it seems okay if you need an answer to a question, but plain bad if I'm looking for a specific page.
e.g. I'm just now looking for the menu for a local restaurant. "restaurantname menu" in Kagi (Google would presumably be similar) returns a link to the menu as the first result in about a second. Or "restaurantname menu !" goes directly to the menu in about a second.
Meanwhile, searching "restaurantname menu" in chatgpt takes about 5 seconds to return an embedded map from mapbox showing the location of the restaurant. If I click the restaurant pin on the map, there's no menu link, the 667 reviews have no link or way to view, and the restaurant description literally says "I don't have enough information to identify which local business <restaurantname> refers to."
Below the map there's some text: "If you mean <restaurantname> in <place>, here’s the current menu. <restaurantname>". The <restaurantname> link just opens the same card as clicking the pin on the map.
After that there's a bullet point list of the menu that ommits a ton of detail and options.
After that there's finally a link... that I can click to open up a popup at the bottom of the page with an actual link to the menu.
This was literally the first thing that popped into my head, I didn't have to put any effort into finding a query where chatgpt falls on its face.
Yeah, I would... but also a search engine works even better than a map if it's a restaurant I'm already familiar with and just want a link to restaurantwebsite.com/dinner-menu
Yeah seems pretty obvious to me that most people are not going to be using a search engine in 5 years. In the sense of searching for something and combing through the results to find the answer.
Divide et impera. People need to learn. Get together, do things together, create alternatives, use that, maintain position, do not let outsider saboteurs neuter the project. Problem solved.
"Figuring it out" doesn't help if it's encrypted with a key only Google holds or a reference to some database record that you don't have. In either case, the only way to resolve this is to ask Google, which means they can track it as a click (and rate limit etc.).
Note: Article published by Autom.dev which, from a quick read of their homepage, seems like it scrapes Google search results in violation of Google's terms of service and sells those results to customers via an API. That's just my quick read of it, though, so this could be wrong.
Anecdata: I've observed this behavior for several weeks already as a regular user without a Google account and there are countless comments from regular users reporting the same behavior.
It's relevant to know when the source of information is biased due to a conflict of interest. The source's business is apparently scraping Google (and Bing, etc).
Google wants to build up a private web here. We already know this from AMP before.
Also, the search results are total garbage now. Just give it a try and you see how
useless the UI is, with AI results first, then tons of commercial go-to entries and
only then a few links that are often also totally useless and irrelevant. Google optimised towards crap.
We really need to get rid of Google. Qwant results are now a bit better
than before, but still not that great, and I hate that its default UI is
a clone of Google. We need more alternatives. Let's get rid of Google once
and for all - it has disappointed too many people now.
now is time to start the ai to spam/reverse the - google.com/goto for url resolution, it can't be that hard to reverse engineer how the hash works. It was doable for twitter and YouTube, it will be doable here too...
why is this such a bad thing? it's not really any different from using a uuid as a user facing key, which basically everyone does.
and trying to protect your moat isn't automatically a bad thing. they clearly feel it's helping competition, so they're closing a hole. competition is good doesn't mean help your competitors.
- coming from someone who's been using fastmail as my personal for ~10 years because i don't want my emails to be backprop fodder
I guess someone made a website which google crawled and adding a senf made uuid to it is like google trying to own it rather than just being a true search engine just having index to it.
ISPs can presumably correlate the Google query string with the request following the response to the goto and so make a search index? I guess they would charge too much.
Do any large ISPs use visit data to feed into a search index?
Unless you're using DNS over HTTPS they can see the unencrypted DNS traffic.
There's also Encrypted Client Hello, but they can also see which IP you're connecting to.
Also if the domain responds on the bare IP or there's nothing else they've seen on that IP it's a fair assumption. This is without even wondering if their router sends telemetry
It hasn’t provided direct URLs for decades? Not exactly new behaviour.
I’ve got something that will blow your mind. Google now has tracking analytics, for get this, your business’s phone number. Some “Adsense partner” convinced our web admin to install a little script which changes your phone number on your website so they track phone call enquires back to search engine leads / advertising spend.
Yeah no thanks, that was creepy as hell and had it rolled back ASAP. You’ve got to realise the power these tech companies hold over your business. Don’t show up in the search results, someone lists your business as closed in maps, a tracking phone number goes dead so they can’t call you, you might as well have shut up shop and ceased to exist.
Would that company then sell that data back to Google? The consolidation of control of information under a single actor is certainly a factor, but it's not that simple, and it's not the only factor. You shouldn't be so eager to outsource as much of your business intelligence as possible.
Previously the clear target URL of a search result was added to the result-link (or div?) as an additional data-tag (aria or whatever, can't remember), and thus was able to be used with a userscript or similar to replace the Googleified tracking redirect link. And it was also part of the googliefied tracking URL, either way you were able to see the clear target URL.
That no longer is possible, you must ask google.com for consent to visit the result link before seeing what the final target URL is.
But the extension doesn't need to use that exact mechanism, and the person that linked it didn't mention specific mechanisms. The purpose of the extension is putting the real URLs back, and it can still do that.
And it can still prevent google from knowing which search results you click on, even though OP didn't mention that feature.
You do a google search. The extension resolves every link on the page immediately via the goto urls. This doesn't leak any information to google because they obviously know which links they sent you. Now the links are resolved and you can copy and click them and get clean URLs, without sending any information about which ones you're copying or clicking.
I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". It is surprising to see how much they have stripped from our view - long tail results, actual results for product reviews and not ad spam, no preference for 20 page recipe sites.
There are still illegal streaming sports and movie sites everywhere (who knew) and all other seedy corners of the internet that have been neatly erased by Google. It makes me nostalgic for that brief window of time when the web was truly uncontrolled, when page rank had meaning and you didn't know if your search would return 0 results or 4,000 pages, which you could actually browse.
reply