Estimating top N hosts in cardinality using small memory resources

Keisuke Ishibashi, Tatsuya Mori, Ryoichi Kawahara, Yutaka Hirokawa, Atsushi Kobayashi, Kimihiro Yamamoto, Hitoaki Sakamoto

Research output: Chapter in Book/Report/Conference proceedingConference contribution

6 Citations (Scopus)

Abstract

We propose a method to find N hosts that have the N highest cardinalities, where cardinality is the number of distinct items such as the number of flows, ports, or peer hosts. The method also estimates their cardinalities. While existing algorithms to find the top N frequent items can be directly applied to find N hosts that send the N largest numbers of packets through packet data stream, finding hosts that have the N highest cardinalities requires tables of previously seen items for each host to check whether an item of an arrival packet is new, which requires a lot of memory. Even if we use the existing cardinality estimation methods, we still need to have cardinality information about each host. In this paper, we use the property of cardinality estimation, in which the cardinality of intersections of multiple data sets can be estimated with cardinality information of each data set. Using the property, we propose an algorithm that does not need to maintain tables for each host, but only for partitioned addresses of a host and estimate the cardinality of a host as the intersection of cardinalities of partitioned addresses. We also propose a method to find top N hosts in cardinalities which is to be monitored to detect anomalous behavior in networks. We evaluate our algorithm through actual backbone traffic data. While the estimation accuracy of our scheme degrades for small cardinalities, as for the top 100 hosts, the accuracy of our algorithm with 4, 096 tables is almost the same as having tables of every hosts.

Original languageEnglish
Title of host publicationICDEW 2006 - Proceedings of the 22nd International Conference on Data Engineering Workshops
EditorsXiaofang Zhou, Roger S. Barga
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)0769525717, 9780769525716
DOIs
Publication statusPublished - 2006 Jan 1
Externally publishedYes
Event22nd International Conference on Data Engineering Workshops, ICDEW 2006 - Atlanta, United States
Duration: 2006 Apr 32006 Apr 7

Publication series

NameICDEW 2006 - Proceedings of the 22nd International Conference on Data Engineering Workshops

Other

Other22nd International Conference on Data Engineering Workshops, ICDEW 2006
CountryUnited States
CityAtlanta
Period06/4/306/4/7

ASJC Scopus subject areas

  • Information Systems
  • Computer Networks and Communications
  • Information Systems and Management

Fingerprint Dive into the research topics of 'Estimating top N hosts in cardinality using small memory resources'. Together they form a unique fingerprint.

  • Cite this

    Ishibashi, K., Mori, T., Kawahara, R., Hirokawa, Y., Kobayashi, A., Yamamoto, K., & Sakamoto, H. (2006). Estimating top N hosts in cardinality using small memory resources. In X. Zhou, & R. S. Barga (Eds.), ICDEW 2006 - Proceedings of the 22nd International Conference on Data Engineering Workshops [1623824] (ICDEW 2006 - Proceedings of the 22nd International Conference on Data Engineering Workshops). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/ICDEW.2006.56