Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Streamingdatenstrukturen zum Analysieren von Nu...
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
Torsten Bøgh Köster
November 03, 2015
Technology
920
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Streamingdatenstrukturen zum Analysieren von Nutzeraktionen in Echtzeit
Torsten Bøgh Köster
November 03, 2015
More Decks by Torsten Bøgh Köster
See All by Torsten Bøgh Köster
LLMs im Griff: Observability, Tracing und Security
tboeghk
0
74
LLMs im Griff: Observability, Tracing und Security
tboeghk
0
59
Oder mache ich es lieber selbst? Wie sich Kosten und Geopolitik auf Cloud-Betrieb auswirken
tboeghk
0
91
Taking an abandoned Solr search from zero to GenAI hero
tboeghk
0
73
Oder mache ich es lieber selbst? Wie sich Kosten und Geopolitik auf Cloud-Betrieb auswirken
tboeghk
0
65
🔪 How we cut our AWS costs in half
tboeghk
0
430
Shared Nothing Logging Infrastructure
tboeghk
0
130
Beyond Cloud: A road trip into AWS and back to bare metal
tboeghk
1
120
Shared Nothing Logging Infrastructure
tboeghk
0
1.4k
Other Decks in Technology
See All in Technology
All About Sansan – for New Global Engineers
sansan33
PRO
1
1.5k
『自分で判断できるか』を基準に、プロダクトのハンズオン研修でAI利用の線を引いてみた / Where We Drew the Line on AI in Hands-on Training
honyanya
0
180
AI時代のアウトプット――変わったこと、変わらないこと / Devsumi 2026 Kansai #devsumi
jnchito
0
590
AWSとAzureのマルチクラウド活用における強い味方___AWS_Kiroを使った二刀流スキル作成.pdf
duelist2020jp
0
110
あなたの知らないバージョン命名規則
sat
PRO
2
810
AI駆動開発を組織で促すために
lycorptech_jp
PRO
4
6.3k
markdown-poster Introduction
kazamori
0
220
[ホンマでっか SRE] あなたはなぜ SRE に?
_awache
0
270
AIレビュー時代に必要なのは、SLOで引く撤退ライン
nobuoooo
0
120
1.5時間を無駄にして学んだwsl2におけるaptとsnapの選択と仕組み
yosaka0123
0
450
Bill One 開発エンジニア 紹介資料
sansan33
PRO
7
20k
マイナンバーカード本人確認の実装比較(OAuth/OIDC Numa (Immersion) Workshop 2026) / 20260825 numa-12
oidfj
PRO
0
300
Featured
See All Featured
Fantastic passwords and where to find them - at NoRuKo
philnash
52
3.8k
Helping Users Find Their Own Way: Creating Modern Search Experiences
danielanewman
31
3.3k
Designing for humans not robots
tammielis
254
26k
Organizational Design Perspectives: An Ontology of Organizational Design Elements
kimpetersen
PRO
1
800
XXLCSS - How to scale CSS and keep your sanity
sugarenia
249
1.3M
Effective software design: The role of men in debugging patriarchy in IT @ Voxxed Days AMS
baasie
0
480
Exploring the relationship between traditional SERPs and Gen AI search
raygrieselhuber
PRO
2
4.2k
Build your cross-platform service in a week with App Engine
jlugia
234
19k
We Analyzed 250 Million AI Search Results: Here's What I Found
joshbly
1
1.9k
Odyssey Design
rkendrick25
PRO
2
780
Into the Great Unknown - MozCon
thekraken
41
2.7k
Marketing to machines
jonoalderson
1
5.7k
Transcript
Streamingdatenstrukturen zum Analysieren von Nutzeraktionen in Echtzeit WJAX 2015 Torsten
Bøgh Köster (Shopping24) 3. November 2015
Agenda
@tboeghk CTO shopping24 internet group Search Technology Meetup Hamburg Search,
build, delivery, code quality, road bike
None
Open Source Power. Delivered.
search system @ shopping24
Anwendungsfall 1
Produkte gezielt löschen
Bloomfilter
Funktionsweise Bloomfilter
Lokale Bloomfilter
Anwendungsfall 2
Relevante Produkte je Suchanfrage
Benutzeraktionen einfangen
Benutzeraktionen verarbeiten
None
You cannot scale into real time!
Stream Mining
Logstash FTW!
Popularitätswerte als Rankingfaktor
Mit Mandanten exponentielle Datenpunkte
None
The Count-Min-Sketch: A Bloomfilter on Steroids
Wie geht das?
None
Relevanz von Datenpunkten im zeitlichen Verlauf
Exponential Decay
Punisher.java
Anwendungsfall 3
Populäre Suchen in der Autocompletion boosten
Heavy Hitters a.k.a. TopK
Und so geht’s
Und sonst so?
BitSet / SparseFixedBitSet: Non-probabilistic existence test
HyperLogLog: Estimating cardinality
Data Sampling: Reduce large data sets using statistics. Use for:
expensive computations
Data Sampling: Existence computation reduces large data sets to constant
~700pcs
Packed Ints: Reduce heap size for large integer arrays
Packed Ints: Further heap reduction with an offset
None
@see ‣T-Digest (Ted Dunning): https:// www.mapr.com/blog/better-anomaly- detection-t-digest-whiteboard- walkthrough ‣Realtime personalization
(Mikio Braun): http:// blog.mikiobraun.de/2014/05/bbuzz- realtime-personalization- recommendation-stream-mining.html ‣Algorithms and data structures that power Lucene and Solr (Adrien Grand): http:// berlinbuzzwords.de/session/ algorithms-and-data-structures- power-lucene-and-elasticsearch ‣HypeLogLog in Reds: http:// redis.io/commands#hyperloglog ‣Sketching & Scaling Series: http://blog.kiip.me/engineering/ sketching-scaling-part-1-what- the-is-sketching/
Questions? @tboeghk developer.s24.com
[email protected]