web crawling, and how does it work • Which types of organisations can benefit from web crawled data and how • The role of web crawling in the rapidly evolving world of AI • How AI has transformed this space completely
collect data from multiple websites • Web crawling is the collection step, indexing makes it searchable • ‘Like someone flipping through all the books in the library, and leaving notes for the librarian’ • Starts at one URL, follows links to discover more • Crawlers use algorithms and obey rules so they don’t overwhelm sites or crawl private areas • Powers search engines, web archiving, content analysis
~160 million actively used • ~50 million company websites • 90% of those are small businesses • ~1 to 1.5 billion subdomains • ~400 trillion pages total • ~50 to 400 billion pages indexed
websites selling their goods or scamming • Bad actors create multiple websites selling counterfeit goods or scam the buyer • AI has rapidly increased the number of these types of stores on the web On Brand Protection
to existing trademarks • Exploit the brand for phishing or scamming, divert web traffic, or try to sell the domain back to the trademark owner • Structured web crawled data in combination with data science can find these domains Second example
website builders • Since the rise of AI in 2023, regular website builders made way for their AI substitutes • Hedge funds are monitoring these developments closely, and invest in early stages of the growth • Structured web crawled data can look under the hood at instalments For Hedge Funds
• Manual work still needed to combine results • Reliance on unstructured data reduces accuracy • Outcomes are often ‘best guesses’ rather than precise insights • Time-consuming to process at scale • Asking the right questions, drawing the right conclusions
learn how people write and communicate • Models learn about factual knowledge or events that are happening • Models learn about perspectives of people • We’re in the middle of the race between LLMs towards the highest quality data • Google’s AI summaries • X’s Grok • OpenAI & Reddit • Meta & illegally obtained ebooks
can expose structured sources, like web crawled data, directly to AI • 'Librarian’s assistant, who gathers insights from organised books, hands them to the LLM for reasoning, brings back clear results to the reader’ • You can ‘talk to the data’ and get human-like responses • AI turns raw numbers into clear summaries, charts, and takeaways • Less manual combining, faster decision-making
produces huge amounts of data • On the right you see a subset of web crawled data, classified under a few random data points • AI now helps to make conclusions and visuals of the already existing structured web data
multiple industries • Crawled data is a major training source for LLMs • Heavy reliance on unstructured data can reduce accuracy • AI + structured data unlocks more reliable and actionable insights