Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Latency could be artificial, in order to get:

-differential pricing

-ability to switch to transparently switch to slower technologies in future



If they're aggressive enough, I'd say there could justifiably be a quite large latency. Say for example their block size is 55 disks with 10% redundancy. Then the seek time is the maximum seek time of the lowest 50 seek times within the set of disks which haven't failed. Now there's a queue to read each drive, and it may very well frequently be a huge maximum queue pretty much every time. Even if the queue is typically short, they would still have to show a typical high value, which I'd say is essentially read time of an entire disk.

Now factor that the disks are both large and crappy. A 120 Mb/s read speed and 1Tb disk size would imply 8000 secs ~ 2 hours. Factor possibility of differential pricing as you mentioned (even longer queues), and you may get an upper bound of 3 or 4 hours.

I'm just speculating though.


S3 has to handle random access patterns, which are the hardest to optimize.

I wouldn't be surprised if Glacier's latency wasn't purely artificial so much as it was a deliberate design decision so the architecture can be very different from S3: pure streaming I/O, huge block sizes, concurrent access is nowhere near the same, etc. That much time allows really aggressive disk scheduling and it'd make it much easier to do things like spread data across a large number of devices with wide geographic separation.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: