Joint work with: Oktie Hassanzadeh, Yang Yang, Jiang Du, Minghua Zhao, Renee J. Miller University of Toronto Christian Fritz University of Southern California
publications pages ¤ Scientists maintain a bibtex file; BibBase does the rest ¤ Publishes them in HTML ¤ Publishes them in RDF ¤ Links entries to the open linked data cloud ¤ With incentive, scientists are helping us build a bibliographic database (think DBLP but automated) ¤ Invaluable data set for benchmarking duplicate detection and semantic link discovery systems
As of yesterday (September 1, 2010) ¤ ~ 100 active users ¤ 4520 publications, 4883 authors, 502 journals, 1881 proceedings, 88 keywords ¤ 39201 author links, 2768 publication links, 30 keyword links ¤ Note that this is before we do any form of “marketing”
“R. J. Miller” or “RJ Miller” ¤ Publication entries ¤ Journal & conferences: “VLDB” or “Very Large Data Base” ¤ Solutions ¤ Local detection (within a single bibtex file) ¤ Global detection (across multiple files)
duplicates. ¤ E.g. within a single file, it is highly likely that “Renee J Miller” is the same as “RJ Miller”. ¤ Users can specify a suffix to the name to differentiate them (DBLP approach). ¤ E.g. “Min Wang” vs “Min Wang2”
record linkage, or reference reconciliation is a well- studied problem and an active research area. [Tutorial- VLDB’05, Tutorial-SIGMOD’06] ¤ We use existing declarative techniques [D.App.σ-SIGMOD’07] to detect duplicates across multiple files. ¤ Display disambiguation page on HTML interface and rdfs:seeAlso attribute on RDF interface. ¤ Also enables user to provide feedback by @string{vldb = Very Large Data Base}
Dynamic grouping of entities (by year, keyword, etc) ¤ RSS feed for notification ¤ DBLP scraper to generate bibtex files from DBLP records ¤ Statistics on usage ¤ Enhancement to existing MIT bibtex ontology file
bibliographic data ¤ Semantic web technologies as a result of complex triplification performed inside the system ¤ Invaluable data set ¤ Future Work ¤ More comprehensive duplicate detection ¤ Links to more external data sources ¤ Better engineering and service level agreement (99.99%?) ¤ Broader user base