Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Streamingdatenstrukturen zum Analysieren von Nu...
Search
Torsten Bøgh Köster
November 03, 2015
Technology
920
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Streamingdatenstrukturen zum Analysieren von Nutzeraktionen in Echtzeit
Torsten Bøgh Köster
November 03, 2015
More Decks by Torsten Bøgh Köster
See All by Torsten Bøgh Köster
LLMs im Griff: Observability, Tracing und Security
tboeghk
0
83
LLMs im Griff: Observability, Tracing und Security
tboeghk
0
62
Oder mache ich es lieber selbst? Wie sich Kosten und Geopolitik auf Cloud-Betrieb auswirken
tboeghk
0
110
Taking an abandoned Solr search from zero to GenAI hero
tboeghk
0
80
Oder mache ich es lieber selbst? Wie sich Kosten und Geopolitik auf Cloud-Betrieb auswirken
tboeghk
0
66
🔪 How we cut our AWS costs in half
tboeghk
0
460
Shared Nothing Logging Infrastructure
tboeghk
0
140
Beyond Cloud: A road trip into AWS and back to bare metal
tboeghk
1
130
Shared Nothing Logging Infrastructure
tboeghk
0
1.5k
Other Decks in Technology
See All in Technology
Issue 駆動でスペシャリストの意図を届ける、AI 実装のアクセシビリティ向上
thkt
0
120
バイブコーディング時代のWebアプリ開発入門~Cloud Runで学ぶセキュアなビルドとデプロイ
waiwai2111
1
130
【技術的負債conf】事業成長に伴う技術的負債の説明責任とAIによるモニタリング、認知的負債について
i35_267
3
1.7k
エージェントはローカル、検証はMicroVM — Lambda MicroVMsでつくるServerless CI
fujioka6789
3
240
AIエージェントを最高のパートナーに育てる方法|評価と判断軸を育てる5つのステップ
koichiaoki
1
150
Claude Codeを「使うほど育つ」AI秘書にするノウハウ
minorun365
PRO
31
27k
事業課題から技術的負債に向き合う
sansantech
PRO
2
1.9k
Genieを崇めよ
kameitomohiro
0
160
作って終わりじゃないサーバーレス 〜9年運用する大規模EC物流API基盤の設計・運用のリアル〜
zozotech
PRO
0
130
LTのテーマ どうきめてる?〜5つの型と私のやり方〜
yama3133
1
110
アプリログインとWeb認証基盤をつなぐ ASWebAuthenticationSession 作法
shimastripe
1
340
Railsのように考える: See through the Master
snoozer05
PRO
4
1k
Featured
See All Featured
XXLCSS - How to scale CSS and keep your sanity
sugarenia
250
1.3M
CSS Pre-Processors: Stylus, Less & Sass
bermonpainter
360
30k
Exploring the relationship between traditional SERPs and Gen AI search
raygrieselhuber
PRO
3
4.3k
The Cost Of JavaScript in 2023
addyosmani
55
10k
Avoiding the “Bad Training, Faster” Trap in the Age of AI
tmiket
0
240
16th Malabo Montpellier Forum Presentation
akademiya2063
PRO
0
380
GraphQLの誤解/rethinking-graphql
sonatard
75
12k
How to build a perfect <img>
jonoalderson
1
6k
First, design no harm
axbom
PRO
2
1.3k
How GitHub (no longer) Works
holman
316
150k
30 Presentation Tips
portentint
PRO
1
400
Practical Tips for Bootstrapping Information Extraction Pipelines
honnibal
25
2.1k
Transcript
Streamingdatenstrukturen zum Analysieren von Nutzeraktionen in Echtzeit WJAX 2015 Torsten
Bøgh Köster (Shopping24) 3. November 2015
Agenda
@tboeghk CTO shopping24 internet group Search Technology Meetup Hamburg Search,
build, delivery, code quality, road bike
None
Open Source Power. Delivered.
search system @ shopping24
Anwendungsfall 1
Produkte gezielt löschen
Bloomfilter
Funktionsweise Bloomfilter
Lokale Bloomfilter
Anwendungsfall 2
Relevante Produkte je Suchanfrage
Benutzeraktionen einfangen
Benutzeraktionen verarbeiten
None
You cannot scale into real time!
Stream Mining
Logstash FTW!
Popularitätswerte als Rankingfaktor
Mit Mandanten exponentielle Datenpunkte
None
The Count-Min-Sketch: A Bloomfilter on Steroids
Wie geht das?
None
Relevanz von Datenpunkten im zeitlichen Verlauf
Exponential Decay
Punisher.java
Anwendungsfall 3
Populäre Suchen in der Autocompletion boosten
Heavy Hitters a.k.a. TopK
Und so geht’s
Und sonst so?
BitSet / SparseFixedBitSet: Non-probabilistic existence test
HyperLogLog: Estimating cardinality
Data Sampling: Reduce large data sets using statistics. Use for:
expensive computations
Data Sampling: Existence computation reduces large data sets to constant
~700pcs
Packed Ints: Reduce heap size for large integer arrays
Packed Ints: Further heap reduction with an offset
None
@see ‣T-Digest (Ted Dunning): https:// www.mapr.com/blog/better-anomaly- detection-t-digest-whiteboard- walkthrough ‣Realtime personalization
(Mikio Braun): http:// blog.mikiobraun.de/2014/05/bbuzz- realtime-personalization- recommendation-stream-mining.html ‣Algorithms and data structures that power Lucene and Solr (Adrien Grand): http:// berlinbuzzwords.de/session/ algorithms-and-data-structures- power-lucene-and-elasticsearch ‣HypeLogLog in Reds: http:// redis.io/commands#hyperloglog ‣Sketching & Scaling Series: http://blog.kiip.me/engineering/ sketching-scaling-part-1-what- the-is-sketching/
Questions? @tboeghk developer.s24.com
[email protected]