Economical and Efficient Big Data Sharing with i-Cloud Thepparit Banditwattanawong Masawee Masdisornchote
Putchong Uthayopas Computer Engineering Department Kasetsart University Bangkok, Thailand
arXiv:2609.23438v1 [cs.DC] 20 Sep 2026
School of Information Technology, Sripatum University Bangkok, Thailand
Abstract—Big data can be hosted on cloud and being shared distributedly through cloud services in an unprecedented volume, variety and velocity. This causes not only cloud network congestions and delayed cloud services but also increases in public cloud data-out charges. Client-side cloud cache alleviates these problems. Furthermore, cloud cache must be aware of nonuniform data-out costs when big data is stored in hybrid clouds built with different public cloud providers. Deploying iCloud approach as the core mechanism of cloud cache could save data-out cost up to 14.78% or 4,425 USD saved per annum based on our representative scenario, and delivered 17.24% bytehit, 17.96% delay-saving and 29.33% cache hit outperforming LRU, GDSF and LFU-DA approaches. A main finding is that iCloud, learning uniform cost patterns, could perform well against nonuniform cost environment. Keywords—Big data, cloud computing, hybrid cloud, cloud cache, artificial neural network, cost-saving ratio.
I.
I NTRODUCTION
Big data such as social media contents, archive of highdefinition videos gathered via ubiquitous information-sensing devices and scientific data could be acquired and stored within an organization’s external cloud(s) and distributedly retrieved by staffs or customers via cloud services offered by the organization. This leads to the downstream bandwidth saturation of network connection between external cloud and big data consumer premise, long-delayed cloud service responsiveness and importantly increases in external cloud data-out charge imposed by public cloud provider [1]–[3]. The significance of the last problem could be realized through the following representative scenario (which is also referred to throughout this paper): an enterprise utilizing big data residing in clouds by transferring it through 10 Gbps Metro Ethernet with 25% average downstream bandwidth utilization for 8 work hours a day, and 260 workdays per year requires the total amount of cloud data-out transfer 190.43 TB per month. This data transfer volume can be translated as 29,933 USD per month based on the weighted average cost 0.1535 USD per GB of Google Cloud Storage’s network egress charge in Asia-Pacific region as of September 2013 [2]. The sharings of big data can be conducted in an economical and network-friendly manner by using client-side cloud cache. Client-side cloud caches are located in or nearby user premise in the form of enterprise-level shared cache, personal webbrowser cache or local user-application cache. Fig. 1 demonstrates the deployment scenario of a shared cloud cache where HTTP requests to external hybrid cloud are proxied by a
HTTP request
Public cloud provider 1 Private cloud
Consumer premise HTTP response (Cache Miss)
Cloud cache
Hybrid cloud interconnection HTTP
Pubic cloud provider 2
t
reques
Private cloud
(Cache hit)
(Cache Miss)
HTTP response
The Internet
Fig. 1.
Cloud cache deployment in a hybrid cloud scenario.
cloud cache, which in turn replies with the valid copies of the requested big data objects either from its local cache repository (i.e., cache hits) or by retrieving updated copies from the cloud (i.e, cache misses). Cloud caches inherit the capabilities of traditional forward web caching proxies since cloud data is also delivered by using the same set of HTTP/TCP/IP protocol stacks as in WWW. Unavoidably, the same problem as in web caching proxies also exists in cloud caches that is caching entire remote data in local cache is not economically plausible, thus cache eviction approach is mandatory for cloud caches. When the big-data hosting cloud is a kind of hybrid, which employs different public cloud providers for risk management purpose, different data-out charge rates potentially apply to data-out costs and must be aware of by cache eviction approach for economical performance optimization. II.
R ELATED WORKS
There are numerous cache eviction approaches in present existence such as [4]–[9]. They have been extensively investigated in our previous works [10]. To recap, none of them aims for big data and cloud computing for two main reasons. First, those approaches evict big objects to optimize hit rates rather than byte-hit and delay-saving ratios, crucial to the scalability of cloud-transport infrastructures and the responsiveness of cloud computing services, respectively. Second, they do not support multiple public-cloud data-out charges, thus neither improve cloud consumer-side economy nor support hybrid cloud deployment. i-Cloud approach, originally proposed in [11], extends its prior nonintelligent versions [10], [12], [13] by integrating an artificial neural network (ANN) to automate an algorithmic parameter self-tuning for workload adaptability. Its performances have been studied without comparing with the other well-known approaches and based on the totally uniformcost circumstances of both ANN training and deployment phases.
TABLE I. Feature Total requests Total requested bytes Total unique objects Maximum bytes of total unique objects
C HARACTERISTICS OF PREPROCESSED TRACES . BO
15 days 352,224 2,294,688,191 181,624 1,386,970,321
NY 31 days 639,187 4,149,211,314 323,979 2,262,144,480
This paper presents the new and comparative performance behaviors of i-Cloud and three well-known approaches by emulating a hybrid cloud as a testing environment where economical costs offered by two public cloud providers are nonuniform. The main objective of doing this is to observe the performances of i-Cloud that has learned uniform cost patterns but is deployed against a nonuniform cost environment. A minor objective is to show how much i-Cloud outperforms the other approaches when data-out charge rates are nonuniform. The findings of these observations would convince users of iCloud performances when deploying cloud cache for a single private cloud at the beginning that later evolves to a hybrid cloud according to new business requirements. III.
O RIGINAL DATA SETS
To understand the following sections, it is necessary to clarify our original trace data sets from which the other types of data sets used in our study were derived. The trace data sets were provided by IRCache project [14] in the form of raw HTTP traces representing the object request streams of two separate user-community behaviors: a 31-day BO trace was gathered from a user community in Boulder from 16th August to 15th September 2012, a 31-day NY trace was collected from the other user community in New York from 16th July to 15th August 2012. To emulate cloud computing request streams of the same consumer organization, these raw traces must be adjusted with realistic assumptions. For this reason, each of these traces was preprocessed to extract only the object references of 50 most popular domain names (to emulate the total number of HTTP domains administrated within the first author’s university). Omitting this preprocessing step was translated that a single consumer organization owns an impractically large number of domain names hosted on its own cloud(s). All references representing dynamic object requests were not excluded during preprocessing stage to reflect actual caching performances against all types of requests. (That is why our performance results were not so high as was reported in some of the related works.) We found that the extracted references reflected cloud traffics or big data as containing both cloud service requests to, for instance, Facebook, Youtube, Twitter SaaSes, and Mediafire IaaS, and WWW requests, which were assumed to go to cloud-hosted web servers. To make the preprocessed traces self-contained, additional preprocessings were the removals of unused fields and appending object expiration times as a new field to all object references. The expiration times were figured out based on the following heuristic rules with an assumption that object creators set object expiration times deliberately: •
Rule 1: by scanning down all references within a trace in timestamp order, object expired immediately after
15 days 639,199 6,499,655,874 290,851 4,791,008,825
31 days 1,311,880 17,067,821,671 593,365 10,801,010,237
its size was found changed from its last reference. •
Rule 2: as long as the size of object remained unchanged, its lifetime was extended to its last request as appeared in a trace.
•
Rule 3: object apparent only once throughout a trace expired suddenly after its use. (This object tends to be a kind of dynamic one, which is by default not cached by Squid. Thus, this rule is prescribed so.)
Once the one-month BO and NY traces had been preprocessed, the references belonging to the first 15 days of each trace were duplicated into a new trace resulting in a 15day BO trace and a 15-day NY trace. This allowed objects referenced in both 15-day traces to have maximum one-month lifespans (as a result of applying the heuristic rule to the onemonth traces) rather than merely 15 days, and avoided false positives for static objects made by the heuristic rule 3. The final preprocessing results are characterized in Table I. Notice that the last feature must be denoted as the maximum bytes (of total unique objects) since the sizes of several unique objects enlarged over a reference stream as detected in each trace. IV.
E CONOMICAL AND T ECHNICAL P ERFORMANCE M EASUREMENT
The algorithm design and results of i-Cloud have been described in standard performance metrics [4], defined as follows. For an Pn object Pn s h / i, byte−hit ratio = P i i i=1 si ri , i=1 P n delay−savingP ratio P = li hi / ni=1 li ri and i=1 n n hit rate = i=1 hi / i=1 ri where si is the size of i, hi is how many times a valid copy of i is fetched from cache, ri is the total number of requests to i, and li is the loading latency of i from cloud. In addition, cost-saving ratio [10] was also used to capture the economical performances of our studied approaches. The metric measures how much money can be saved by serving the valid copies of requested objects from cache. It is expressed as below. Given an object i, Pn ci si h i cost−saving ratio = Pi=1 n i=1 ci si ri
(1)
where ci is the data-out charge rate or monetary cost for loading i from cloud. This metric is particularly useful for a hybrid cloud or a nonuniform-cost model (Cf. Section V-D). V.
i-C LOUD A PPROACH
The design goals of i-Cloud are cloud economy, scalability and responsiveness that can be realized by optimizing costsaving, byte-hit and delay-saving ratios, respectively. Hit rate
TABLE II.
I -C LOUD ALGORITHM
Table II presents the pseudo code of i-Cloud algorithm in details. The main principle behind the scene of i-Cloud is contemporaneous proximity [15]. When cache eviction is needed, i-Cloud is invoked. It first formulates a cluster of incache least-recently-used (lru) objects as many as instructed either by a window size parameter or a required cache space (which is at least a requested missing object’s size and controlled by Squid’s low and high watermarks [16]) depending on which one is larger, it will be chosen. Once the cluster of lru objects has been formed, i-Cloud quantifies a profit associated with each object inside the cluster as follows: given an object i, prof iti = si · ci · li · fi · T T Li where si is the size of i, ci is data-out charge rate for loading i, li is latency for loading i, fi is the access frequency of i, and TTLi is the remaining lifespan of i. (This profit formula is justified in [10].) An object with least profit is evicted first from cache. This object eviction process is repeated on the next least profitable objects in the cluster until gaining enough cache room. A significant parameter influencing all performance aspects of i-Cloud is window size (ws in Table II). Our experiments indicated that window size varied at least from one workload to another. Seeking an optimal window size is not an easy task. Lessons learned from our previous studies [10], [12], [13]
Profit of k-LRU cached objects
... ...
MLP forcasts a low-level window size.
Distal teacher
has gained declining impact since these days with globally available broadband network infrastructures, it is perceived that loading remote small objects is fast as if they were fetched from user locus whereas the problems still insist on retrieving big objects over the network (more details are available in [10], [12], [13]).
Required cache space
Input vector normalization
algorithm: i-Cloud input variables: cd /*cache database (recency-keyed min-priority queue)*/ ws /*window size*/ rs /*required cache space*/ local variables: ecd /*empty cache database (recency-keyed min-priority queue)*/ oc /*an object cluster of lru objects (profit-keyed min-priority queue of evictable objects)*/ co /*a candidate object to be included in a cluster*/ ts ← 0 /*total size of ws objects initialized to zero*/ eo /*an evicted object*/ c ← 0 /*counter for objects in a cluster initialized to zero*/ begin if cd.getT otalN umberOf Objects() < ws then ws ← cd.getT otalN umberOf Objects(); ecd ← cd; do co ← ecd.removeLeastRecentlyU sedObject(); ts ← ts + co.getSize(); oc.addObject(co); c ← c + 1; while (c < ws) ∨ (ts < rs); do eo ← oc.removeM inP rof itObject(); cd.evict(eo); while cd.getF reeSpace() < rs; return cd.getF reeSpace();
Missing object to be cached
WS
Cloud cache replacement
Forecaster
1st 2nd wsth kth Cloud cache before cache miss resolution
Fig. 2.
Cloud cache after cache miss resolution
Next capacity miss occurs
Conceptual framework of i-Cloud.
indicated that both current cache state, influenced by preceding object requests (i.e., there existed correlations in the request streams) together with allocated cache size, and required cache space were main factors to window size optimization. Therefore, i-Cloud has integrated a multilayer perceptron (MLP) as a forecaster component to automatically learn both of the main factors to forecast near-optimal window sizes to be input into ws in Table II. Fig. 2 demonstrates the conceptual framework of i-Cloud. It consists of two main processing modules, forecaster and cloud cache replacement. They operate as follows. When a capacity miss (i.e., a request for an object that was in a cache but has been since purged, thus cache eviction is required to serve the request) takes place, the vector of k lru cached objects’ profits together with a required cache space are fed into the forecaster to forecast a near-optimal window size. During this stage, each input vector is passed internally into the input vector normalization process then the normalized vector is presented to the MLP component to forecast a low-level window size. Such a low-level window size is later denormalized by the distal teacher component to obtain a practical (and potentially near-optimal) window size. The practical window size is subsequently presented to the cloud cache replacement module, which follows i-Cloud algorithm in Table II. Regarding MLP structure that prescribes the number of layers and nodes to be placed in each layer of the MLP, we aimed at the smallest structure possible since too large MLP learns training set well but is unable to generalize. We have designed the MLP structure based on the following guidelines. •
The appropriate number of input nodes were determined by two lessons learned from our previous studies [10], [12], [13]. First, both current cache state, which is influenced by past object request stream together with total allocated cache size, and required cache space were main factors to window size optimizations. This design guideline has been implemented as shown in Fig. 2: a current cache state is captured in the form of a profit vector while a missing object size is used to indicate a minimum required cache space. Second, the optimal window sizes ranged between 100 and 8,000 meaning that the lowest profitable objects to be purged from cache were potentially found inside a cluster of 8,000 lru cached objects. Leveraging this fact simplifies and practicalizes the MLP’s implementation since a current cache state can be snapshot by examining only
the part of cache database rather than the whole one, which can be much more time consuming. However, to give chance for any unveiled maximum values of optimal window sizes, we simply chose 10,000 (i.e., k in Fig. 2) lru cached objects to represent each current cache state to be examined by the MLP. Thus, the total number of input nodes is 10,002 including a required cache space node and a bias one. •
It is conventionally guided that some continuous functions cannot be approximated accurately by singlehidden-layer MLP whereas two hidden layers are sufficient to approximate any desired bounded continuous function, which is also the case of i-Cloud. Moreover, the hidden layers are usually kept at approximately the same size to ease training.
•
The output layer necessitates a single node to deliver an estimated low-level window size.
Several structures were experimented in an effort based on these guidelines. We finally came out at a minimal structure that can be expressed in the conventional notation of 10,002/2/2/1, which means 10,002 input nodes, two hidden layers with two nodes each, and a single output node. All the nodes are fully-connected between two adjacent layers, except a bias node is connected to all noninput ones. A complete input vector presented to the MLP is denoted as X =< 1, p1 , p2 , ..., p10,000 , rs >
(2)
where pi=1 to 10,000 is the profit of ith lru object and rs is a required cache space. Notice that the bias is set to a constant activation 1. Every link connecting node j to node i has an associated weight wij .
TABLE III.
F ORECASTER ALGORITHM
algorithm: Forecaster input variables: cd /*cache database (recency-keyed min-priority queue)*/; rs /*required cache space*/; wij /*inter-node connection weight*/; local variables: ecd /*empty cache database*/; ip /*MLP’s input node*/; i, j /*node indices where j is on preceding layer of i*/; temp /*temporary object*/; l /*layer index*/; ui /*weighted sum*/; yi /*output activation*/; begin 1...10000
ecd ←− cd; /* duplicating 10,000 lru cached objects */ /* Read a current cache state into 10,002 input nodes */ ip0 ← 1; /* bias constant */ for j=1 to 10000 do temp ← ecd.removeLeastRecentlyU sedObject(); ipj ← temp.getP rof it(); ip10001 ← rs; /* Normalize inputs except bias node */ for j=1 to 10001 do ipj ← 2 ipj / 1014 − 1; /* Propagate the inputs forward to compute an output */ for each noninput layer ℓ do for each node i in ℓ do P w ip ; ui ← j ij j yi ← sigmoid(ui ); /* Denormalize the output */ y10006 ← 10000 y10006 ; return y10006 ; /*return window size*/
A. Learning phase Besides the MLP structure, to obtain the complete MLP mandates the appropriate set of weights on all node-connecting links. This has been done by means of supervised learning with a distal teacher and back-propagation. We used mean squared error as an objective function. Our MLP learns patterns inherent in caching state history. Each input pattern is organized into a training vector of the form X (Eq.(2)) and was generated every time capacity miss occurs during the i-Cloud simulation of a certain trace, cache size and respective (bytehit or uniform-cost-saving) optimal window size. The result of each complete simulation session was a sequence of input patterns, which were contained inside a single training data set. We generated two distinct data sets by the separate simulation sessions of two 15-day traces (Table I) using 10% cache size (i.e., percent of the maximum bytes of total unique objects referenced in each trace). The first data set, BO15D10%, contained totally 54,639 patterns, while the other data set, NY15D10%, contained 62,607 patterns. We normalized every element value within each input vector except the bias constant by max-min linear scaling. Because the MLP outputted a lowlevel window size of domain [0.0, 1.0] via sigmoid function, to achieve desired window sizes requires a distal teacher to denormalize the low-level window size to obtain an actual window size of [0, 10,000]. Target window sizes were the static optimal ones aforementioned. Note that any actual window size near the lower bound tended to be increased later by i-Cloud
as coded in Table II to be able to fit a required cache space. We used adaptive learning rate instead of fixed one to make the forecaster converge quickly. The learning phase terminated when mean squared error stabilized below 1000. B. Validation phase Our previous work [11] has shown that measuring the accuracy of the trained forecaster by using validation data sets to measure errors produced by the forecaster was almost meaningless as a local optimum window size could have the high degree of validation error. Instead, we have evaluated the performances of i-Cloud including the trained forecaster as a whole by trace-driven simulations in the four performance metrics as described in Section VI. Table III presents the forecaster’s pseudo code, basically comprising an input vector normalization, a feedforward propagation and an output denormalization (i.e., a distal teacher) parts, respectively. To become an effective forecaster, it is mandatory to train the MLP as described in the following subsection. C. Algorithmic practicality Because the total number of cached objects (N) can be increasingly large at runtime. It is necessary to conduct the
0.294
0.140
0.293
LRU
0.137
0.291
GDSF
0.136 0.129
0.160 0.155
0.122
0.150
0.115
0.145 10
20
10
30
0.135
LFUDA 0.290 i-Cloud (BO15D10%)
0.133
0.288
i-Cloud (NY15D10%)
0.130
0.287
Infinite cache size
0.128
0.285
30
10
20
30
10
20
30
Cache size (%)
The comparative performance results of i-Cloud and various approaches based on the 31-day BO trace. 0.181
Byte-hit ratio
0.072
0.054
0.072
0.054
0.036
0.036
0.018
0.018 10
20
30
10
20
0.208
0.177
LRU
0.203
GDSF 0.173
Hit rate
0.090
0.090
Fig. 4.
20
Delay-saving ratio
Fig. 3.
Cost-saving ratio
0.165
Hit rate
0.142
0.170
Delay-saving ratio
0.175
0.143
Byte-hit ratio
Cost-saving ratio
0.150
0.169
0.198
LFUDA
i-Cloud (BO15D10%)
0.193
i-Cloud (NY15D10%)
0.165
0.188
0.161
0.183
30
10
20
30
Infinite cache size 10
20
30
Cache size (%)
The comparative performance results of i-Cloud and various approaches based on the 31-day NY trace.
algorithmic practicality analysis of both modules in Fig. 2. As for the time complexity analysis of i-Cloud algorithm (Table II), the statements that take significant part in processing time are: replicating cd into ecd takes O(NlogN); the first do − while loop takes O(NlogN) as the window size can be set via the preceding if − then statement to as many as N and removing each object from ecd is O(logN) just like adding each object into oc; the second do − while loop has the worstcase running time of O(NlogN) because the number of evicted objects is bounded by N, while removing each object from oc takes O(logN) and deleting an object from cd is O(logN). The other statements are all identically O(1). Therefore, the algorithm is O(NlogN).
rates were associated with the 50 domains of each trace in an interleaving manner so that the nonuniform cost model could be emulated for each trace. Objects retrieved from the same domain were always charged at the same rate.
Regarding the time complexity of the forecaster algorithm (Table III): copying 10,000 lru objects from cd to ecd takes O(logN), the first f or loop takes O(logN) to remove lru object from cd, while the other statements are all identically O(1). Thus, the algorithm is O(logN).
The one-month traces in Table I were used to conduct the separate trace-driven simulations of i-Cloud learning BO15D10 data set, i-Cloud learning NY15D10 data set and the other three well-known approaches, LRU, GDSF, and LFUDA, supported by popular Squid caching proxy [16]. The comparative simulation results in four performance metrics are demonstrated in Fig. 3 and 4. The simulated cache sizes are dictated in percents of the maximum bytes of total unique objects of each of the one-month traces. Interesting observations are as follows.
Since the cloud cache replacement and forecaster modules are sequentially connected as shown in Fig. 2, i-Cloud algorithm is totally O(NlogN) + O(logN) equal to O(NlogN). In other words, i-Cloud strategy can be implemented.
As a remark, although the evaluation relied on the nonuniform cost model, it was our intention to train i-Cloud by using static optimal window sizes, which were tuned based on a uniform cost model rather than nonuniform one. This was to observe how efficiently i-Cloud learning uniform-cost data sets performed against nonuniform-cost ones. VI.
R ESULTS AND DISCUSSIONS
D. Monetary cost models
•
Since this paper aims for multi-provider hybrid clouds, we engaged a nonuniform cost model for evaluation. In nonuniform cost model, all of the 50 domains (Cf. Section III) of each trace were assumed to go to a hybrid cloud, running in two distinct public clouds that offered different data-out charge rates. The first charge rate was set to 0.1535 USD/GB based on Google Cloud Storage (Cf. Section I), while the other was set to 0.0829 USD/GB based on Amazon S3’s data transfer out to Internet charge (US Standard) as of September 2013 [1]. (Both flat rates are actually the weighted averages of the actual regressive rates of Google Cloud Storage and Amazon Simple Storage Service covering the monthly data transfer of the representative scenario described in Section I.) Both flat
It is obvious that i-Cloud of both learning data sets has outperformed all competitive approaches in all simulation cases in not only cost-saving but also byte-hit, delay-saving and hit performance metrics. This means that though i-Cloud learned uniform cost patterns, it has delivered good performances in nonuniform cost environment.
•
i-Cloud performances have shown to stabilize at close degrees to those of infinite cache size in nonuniform cost environment. This is also the strength of i-Cloud against uniform cost models that has been known in our previous works [10]–[13].
•
i-Cloud that learned only small data set derived from 15-day trace has still been able to perform relatively
well against the one month workloads of both user communities. This substantiates the successful longterm deployment of i-Cloud up to some extent. •
Based on the BO trace and cost-saving metric, i-Cloud(BO15D10%) has outperformed iCloud(NY15D10%) at 10% cache size and delivered identical performance to i-Cloud(NY15D10%) at 20% and 30% cache sizes. Using the NY trace, i-Cloud(NY15D10%) has more economized than i-Cloud (BO15D10%) at 10% and 30% cache sizes and less economized than i-Cloud (BO15D10%) at 20% cache size. Hence, in most cases i-Cloud has performed economically better against user community behavior i-Cloud has learned.
•
i-Cloud could retain both high byte-hit and hit rates at the same time based on our simulated workloads. This means that the breakthrough finding of our previous works ( [11]–[13]) that optimal hit and byte-hit ratios could be attained at the same time also applies to nonuniform cost environment. Since i-Cloud tends to evict smaller objects, the finding is an objection to the conventional rule of thumb in traditional web caching: “strategies that tend to remove bigger objects improve the hit-rate but decrease the byte-hit-rate” [4]. VII.
C ONCLUSION
This paper presents i-Cloud cache eviction approach that accommodates the distributed sharings of big data. i-Cloud has access recency as a priority factor for object replacement decision. i-Cloud parameterizes an MLP-based self-tuning window size to generalize the recencies of objects within a formulated object cluster. The lowest profitable clustered objects are purged from cloud cache. Based on the tracedriven simulation results, the distributed sharings of big data was most efficient when employing i-Cloud. Although i-Cloud has been trained based on a uniform cost model, it performed well against a nonuniform cost environment or multi-provider hybrid cloud. ACKNOWLEDGMENT This research is financially supported by Thailand’s Office of the Higher Education Commission, Thailand Research Fund, and Sripatum university (grant MRG5580114). The authors also thank Duane Wessels, National Science Foundation (grants NCR-9616602 and NCR-9521745) and the National Laboratory for Applied Network Research for the trace data used in this study. R EFERENCES [1] [2] [3] [4] [5]
Amazon.com, Inc. (8 August 2013) Amazon simple storage service. [Online]. Available: http://aws.amazon.com/s3/pricing/ Google Inc. (8 August 2013) Google cloud storage. [Online]. Available: https://cloud.google.com/pricing/cloud-storage/ Microsoft. (8 August 2013) Windows azure. [Online]. Available: http://www.windowsazure.com/ S. Podlipnig and L. Böszörmenyi, “A survey of web cache replacement strategies,” ACM Comput. Surv., vol. 35, no. 4, pp. 374–398, Dec. 2003. J. Cobb and H. ElAarag, “Web proxy cache replacement scheme based on back-propagation neural network,” J. Syst. Softw., vol. 81, no. 9, pp. 1539–1558, Sep. 2008.
[6] S. Romano and H. ElAarag, “A neural network proxy cache replacement strategy and its implementation in the squid proxy server,” Neural Comput. Appl., vol. 20, no. 1, pp. 59–78, Feb. 2011. [7] W. Ali and S. M. Shamsuddin, “Intelligent client-side web caching scheme based on least recently used algorithm and neuro-fuzzy system,” in Proceedings of the 6th International Symposium on Neural Networks, 2009. [8] S. Sulaiman, S. Shamsuddin, F. Forkan, and A. Abraham, “Intelligent web caching using neurocomputing and particle swarm optimization algorithm,” in Modeling Simulation, 2008. Second Asia International Conference on, 2008, pp. 642–647. [9] W. Tian, B. Choi, and V. V. Phoha, “An adaptive web cache access predictor using neural network,” in Proceedings of the 15th intl. conf. on Industrial and engineering applications of artificial intelligence and expert systems, 2002. [10] T. Banditwattanawong, “From web cache to cloud cache,” in Advances in Grid and Pervasive Computing, ser. Lecture Notes in Computer Science, R. Li, J. Cao, and J. Bourgeois, Eds. Springer Berlin / Heidelberg, 2012, vol. 7296, pp. 1–15. [11] T. Banditwattanawong and P. Uthayopas, “An intelligent cloud cache replacement scheme,” in Advances in Information Technology, ser. Communications in Computer and Information Science, B. Papasratorn, N. Charoenkitkarn, V. Vanijja, and V. Chongsuphajaisiddhi, Eds., vol. 409. Springer, 2013, pp. 23–34. [12] ——, “Cloud cache replacement policy: new performances and findings,” in Annual PSU Phuket, 2012. PSU PIC 2012. 1st International Conference on, 2013. [13] ——, “Improving cloud scalability, economy and responsiveness with client-side cloud cache,” in Electrical Engineering/Electronics, Computer, Telecommunications and Information Technology, 2013. ECTICON 2013. 10th International Conference on, 2013. [14] National Laboratory for Applied Network Research. (2012) Weekly squid http access logs. [Online]. Available: http://www.ircache.net/ [15] T. Banditwattanawong, S. Hidaka, H. Washizaki, and K. Maruyama, “Optimization of program loading by object class clustering,” IEEJ Transactions on Electrical and Electronic Engineering, vol. 1, no. 4, pp. 397–407, 2006. [16] D. Wessels, Squid : the definitive guide. O Reilly, 2004.