a terminal screen β’ Structured β’ Readable β’ Navigable (search & lookup) β’ Not fancy, but practical β’ ... for people who work on command line tanelpoder.com
N ... BPF_HASH(syscall_id) 10 42 42 42 42 42 BPF_HASH(syscall_ustack) 10 11 42 N ... ... tracepoint:raw_syscalls:sys_enter { @syscall_id[tid] = args->id; } 42 42 42 42 42 42 We are not tracing, logging, appending all events We update, overwrite the current, latest action in custom state arrays ... Populating & sampling the thread state "array"
N ... BPF_HASH(syscall_id) tracepoint:raw_syscalls:sys_enter { @syscall_id[tid] = args->id; } A separate, independent program samples the state arrays using its desired frequency and filter rules to userspace BPF_HASH(syscall_ustack) interval:hz:1 { print(@SAMPLE_TIME); print(@syscall_id); } 10 11 42 N 10 11 42 N 10 11 42 N 10 11 42 N Populating & sampling the thread state "array"
N ... BPF_HASH(syscall_id) tracepoint:raw_syscalls:sys_enter { @syscall_id[tid] = args->id; } BPF_HASH(syscall_ustack) interval:hz:1 { print(@SAMPLE_TIME); print(@syscall_id); } 10 11 42 N 10 11 42 N 10 11 42 N 10 11 42 N The sampler can be an eBPF program (bpftrace, bcc, libbpf) or an userspace agent that reads the maps' pseudofiles Populating & sampling the thread state "array"
every single event to output β’ Unrealistic amount of output & high instrumentation overhead β’ We do not sample only on-CPU threads β’ The profile event only samples on-CPU threads (also commands like perf top by default) β’ We will additionally use the finish_task_switch kprobe for thread sleep (off-CPU) analysis β’ We will "trace" the latest thread state changes into a custom array β’ And "clients" then periodically sample the thread state array & consume the output
Python, etc β’ Currently you get stacks & symbols only for compiled binaries with symbols or debuginfo available β’ It is possible to add higher-level language runtime support and Java runtime-optimized code β’ This has already been done by other tools and works β’ What's the performance overhead? β’ Test it out! J β’ Still beta, I have 6-7 categories of ideas for further improvement β’ It doesn't matter how frequently the frontend samples the TS arrays, doesn't slow others down β’ Will this work with distributed systems? β’ Yes, but not yet implemented (for example, capture + include end-to-end traceID in TS2 array) β’ Distributed systems are still just a bunch of individual systems - that talk to each other β’ Instrumentation is investment! tanelpoder.com
20 or later β’ bcc-tools package installed β’ xcapture-bpf running as root β’ But any Linux user with read access can read its output files! β’ debuginfo in some cases (ideally) β’ xcapture-bpf isn't showing some-detail-I-want (like syscall or IO latencies) β’ I have built out less than 5% of what this method & implementation can provide! β’ BPFapproaches are not only customizable, but completely programmable β’ You can access all kernel events & structures related to thread execution and access userspace memory tanelpoder.com
β’ Evangelize! So that drilldown into thread activity eventually makes sense to everyone! β’ Proper documentation, examples (and a man-page) β’ Optimize the BPF kernel-space performance, also userspace record extraction β’ Profile the instrumentation code itself β’ Improve stack-tracking hashmap to lower memory usage β’ Make some instrumentation dynamic/optional (get_stack on every N iterations) β’ Proper distro packaging β’ Automated CSV compression, archiving (optionally convert to parquet format too) β’ Release v2 GA (September 2024?) β’ For future v3 use libbpf, allow multiple independent samplers of the BPF program maps tanelpoder.com