fix: count sites instead of hostnames in public_hash_list - #357
Conversation
|
This is a follow-up to #324. |
|
TIL |
Signed-off-by: Max Ostapenko <1611259+max-ostapenko@users.noreply.github.com>
…domains and exception rules Signed-off-by: Max Ostapenko <1611259+max-ostapenko@users.noreply.github.com>
|
Hi @tomayac,
Could you please:
WITH private_suffixes AS (
SELECT suffix, is_wildcard
FROM ${ctx.ref('urls', 'public_suffix_list')}
WHERE is_private AND NOT is_exception
),(and remove the sqlStringList helper).
Your core SQL logic for site counting in public_hash_list.js will stay intact, and the suffix list will stay automatically up to date! |
…ines Signed-off-by: Max Ostapenko <1611259+max-ostapenko@users.noreply.github.com>
The >=100 threshold counted distinct hostnames of the resource URL, so a resource used only on many subdomains of one site (for example $city.$vulnerable-group.com) passed the privacy gate. Count distinct sites (eTLD+1 of the embedding page) using the full Public Suffix List, including the private section that BigQuery's NET.PUBLIC_SUFFIX() skips. The private rules come from the urls.public_suffix_list table, which the crawl_complete DAG keeps current.
342ae32 to
80cc0ae
Compare
|
Thanks, @max-ostapenko, that's cleaner. Rebased on |
The ≥100 threshold in
public_hash_listcounts distinct hostnames of the resource URL, so a font used only on 100$city.$vulnerable-group.comlanding pages passes, even though it identifies a single site, and any origin could then probe Cross-Origin Storage for its hash to infer that the user belongs to that group. This PR counts distinct sites (eTLD+1 of the embedding page) using the full Public Suffix List, including the private section that BigQuery'sNET.PUBLIC_SUFFIX()skips, soalice.github.ioandbob.github.iostill count as two sites, renamesnum_originstonum_sites, and adds a script plus a monthly workflow that keep the private suffix rules current. @max-ostapenko, could you please review?