Comment by ivanmontillam
21 hours ago
What I really love about Onion sites is that if they are big enough, performance engineering really becomes Tor-specific. A few examples:
- Making assets embedded as base64 (img src the header logo as base64, all CSS should be inline, etc.).
- Leveraging CSS as much as possible (if you use animations and transitions, use CSS as much as possible for these, avoid JS for them).
- Make sure your website is mostly rendered on the backend. If you're to have JS, your website should work without it.
- Security becomes REALLY fun, as in, avoid XSS, CSRF, SQL Injection attacks and any other injections as much as possible.
As someone summarizes in another comment[0], keep the chattiness as minimal as possible. By chattiness I understand they mean, pack as much data as you can in the same Keep-Alive connection. Avoid making new HTTP requests as much as possible, as each one might get assigned to a new Onion route making things slow.
If you can ship your website to the browser in a single connection, you've won.
I've always been impressed by performance of these big Onion sites, they really push the limits of software engineering creativity, given these constraints and nature of Tor.
--
[0]: https://news.ycombinator.com/item?id=49872320
EDIT: Formatting of bullet points.
Also DDoS becomes a problem, and the ways it's done are pretty specific to Tor. Double captchas are necessary when you're getting DDoSed, due to performance reasons. Oh, and captchas are also pretty specific to Tor as well.
There's also a problem of site fronting. Anyone could run a proxy pretending to be you, for arbitrary reasons (not even necessarily the obvious forging and credential stealing). Every site, even a personal blog, has dozens of parasitic fronts, either actively malicious or dormant. You need off-site ways to tell users what is the real address, and provide a smoke test for them (often a part of the address as a picture, for example in a captcha).
>avoid JS for them
Using any JS defies the point and makes your site instantly suspicious.
> There's also a problem of site fronting. Anyone could run a proxy pretending to be you, for arbitrary reasons (not even necessarily the obvious forging and credential stealing). Every site, even a personal blog, has dozens of parasitic fronts, either actively malicious or dormant. You need off-site ways to tell users what is the real address, and provide a smoke test for them (often a part of the address as a picture, for example in a captcha).
How can you do this without relying on the normal web? Let’s say you use a normal website to show the onion link, if the website gets taken down, you lost your user-trusted mean to do that.
If the website is available on both clearnet and Tor, add a `Onion-Location` header.
https://community.torproject.org/onion-services/advanced/oni...
This way you advertise the onion domain through an established chain of trust and visitors can decide to use that the next time.
>How can you do this without relying on the normal web?
How can you trust anything you haven't experienced personally? By using chains of trust, of course. There are directories that list onion sites, and also sites that link to their peers. That way you can be sure you're still in the same bubble at least, and convert the problem into trusting the entire bubble. It's not automated and pretty ad hoc, if that's what you're wondering. Automation in Tor has a history of being circumvented or exploited with novel scams, this is an adversarial environment.
Your users should have bookmarked it
5 replies →
I've done some searching and can't find a description of "site fronting" that fits with my read of your comment.
I thought the whole point of Tor is that I (and only I) am able to serve traffic at a .onion URL that I have generated. How could someone else get in front of that?
The attacker simply proxies your site on another .onion address and advertises that fake address in a popular directory or an ad. Any visitors coming through that fake URL interact with your site through the attacker's front. Since .onion URLs aren't easy to tell apart, this kind of phishing works well. If you run a discovery bot you can observe that the .onion zone is full of these fake fronts for all kinds of sites, because even if your site aren't of any interest to an attacker but you link to any other site, then the attacker of that site needs to front yours with that link replaced with a fake, to create a separate circle of trust for the victims.
That's why you need to very carefully choose a trusted entry point into the Torosphere and revise your choice from time to time; always remember your initial entry point because if there's any suspicious drama around it or if you notice a mismatching link on different sites, you might have been duped to enter an impersonator-controlled bubble.
Many things can be done about it: claiming your spot in the directories, smoke test captchas, chains of trust, or you can brute force your .onion to find a valid one that starts with a memorable string (the longer the better, but also the harder). Vanity addresses like this deter non-targeted adversaries by requiring proof of work to forge the lookalike URL. None of that is bulletproof, of course.
How are captchas done?
A combination of different ones, sometimes close to what you know from the clearnet, i.e. click pictures that match a description. For the proxy defeating one, the server can send you a picture of the real .onion address where some letters are blanked out, and with noise added to it to prevent it from being solved by the proxy. You then have to enter the blanked out letters from the browser URL.
This can of course still be defeated if the proxy URL receives very few requests and just have a human in the middle to solve the CAPTCHAs. But it does make the attack non-automated.
1 reply →
Are these simply good ideas regardless of tor?
Depends on which ones. Embedding assets as base64 makes little sense nowadays with http pipelining.
Relyinging the least possible on js, and using CSS for animations sounds like good engineering to me
Not really. The latency between a client and clearnet sites is tiny, and splitting assets in to separate resources makes caching work better.
Minimizing latency and connection count is only half of the problem, though. To the greatest extent possible not relying on JavaScript, and to the greatest extent possible not relying on third party libraries when you do seems like a good general takeaway.
Some people live in Australia though
With CDNs of today, they are not so much relevant for the clearnet.
Yes, but
Apart from the 1st one these just seem like good practices in general
Until you got to the network traffic tricks, those were good practices for any landing page you want to open quickly!