As much as I like kagi and wish it success, it's not a search engine from scratch. Kagi uses other search engines (Google and Bing) wraps them and does a light reranking
Kagi is also building their own index at the background, and mixes these indexes as you search.
When I search Kagi for "Hacker News", results start with this fine text:
65 relevant results in 1.09s. 47% unique Kagi results.
So, other indexes are fillers for Kagi's own index. They can't target their bots to places, because they don't have the users' search history. They can only organically grow and process what they indexed.
How is it possible that a search for "Hacker News" produces only 65 results? There are thousands of pages out there with that exact phrase on it (including many sub-pages of this site).
The first result is almost assuredly the right one, but either they're ruling out a lot of pages as not-what-you-meant, or their index is really small.
It's another feature of Kagi. They know they have thousands of results, but they provide you a single page of most relevant results. If you want to see more because you exhausted the page, there's a "more results" button at the bottom.
Kagi reduces mental load by default, and this is a good thing.
That's an interesting positive spin on what is clearly a cost-cutting feature. I'm on the unlimited search tier, so it hasn't really been a big deal, but it's worth noting that clicking the "more results" button charges your account for an additional search too. Or at least it used to.
Key differentiator: Kagi still properly responds to the negation sign and quotes in your search terms. This is why I pay for it. The signal-to-noise ratio is way higher than with other engines.
That page includes the text "Our search results also include anonymized API calls to all major search result providers worldwide".
They source results from lots of places including Google. One way that you can confirm this is to search for something that only appears in a recent Reddit post. Google has done a deal with Reddit that they're the only company allowed to index Reddit since the summer.
edit: I don't think this is a bad thing for Kagi. I'm a very happy subscriber, and it's nice for me that I still get results from Reddit. They're very useful!
Note that due to adversarial interoperability, search engines other than Google can scrape Reddit if they try hard enough. A rotating residential proxy subscription, while pricey, likely still costs orders of magnitude less than what Google paid. The same goes for Stack Overflow. You can also DIY by getting a handful of SIM cards. CGNAT, usually a scourge, works in your favour for this application since Reddit can't tell the difference between you loading 10000 pages and 10000 people on your ISP loading one page each (depending on the ISP)