To continue the theme from last week, backup has traditionally always been a challenge. Backups don’t get completed in time, backup jobs fail, offsite media storage is often haphazard, etc. But there is good news. There are now alternatives in the world of data protection which can significantly reduce dependence on traditional backup & recovery solutions. These alternatives offer many benefits: reduction of the operational interference inherent in long backup jobs, better security, better data availability in the event recovery is needed, data “immutability” and operational cost savings.
One aspect of this new universe of data protection is the concept of simply not backing up the data, at least not in the traditional sense. This is the topic of this post.
First, let’s make a distinction. There are two primary types of data: structured and unstructured. Structured data is simply that which is stored in a structured format. Examples would be databases like SQL Server, MySQL, Oracle, Exchange and the indexes of document management systems. Unstructured data is that which is stored without an identifiable structure, i.e. file data such as MS Office documents, PDFs, Pictures and other image files, movies, etc. Structured and unstructured data tend to have very different usage characteristics. Structured data tends to change rapidly. Think of a business transaction database, such as a point-of-sale or order entry system, data that is changing all through the business day. Unstructured data tends to change slowly. While the overall volume of unstructured data may be increasing rapidly (some analyst estimates are as high as 500% over the next five years) the files themselves change very little or not at all. Unstructured data is typically the largest component of business data. I had a meeting with a client yesterday who said that his unstructured data comprised about 85% of the total in his shop.
There is another important distinction between structured and unstructured data as it relates to backup. Structured data tends to have a higher impact on the business when it is not available. When databases are down, work isn’t getting done. With unstructured data this is generally less of an issue . Therefore, structured data has more stringent RPO and RTO requirements. (RPO, or Recovery Point Objective, is the point in time of the last good copy of the data. If you are backing up nightly, your RPO is 24 hours, not necessarily a good thing if your transaction database fails at 5PM the following afternoon, i.e. a whole day’s transaction data goes up in smoke. RTO is Recovery Time Objective, meaning the length of time required to get a system up and running once it goes down.) Since the two types of data have a different business impact when unavailable, it makes sense that there should be separate data protection approaches for each. It is common, however, for small business to backup everything all at once, which can lead to many of the problems mentioned in the first paragraph.
Unstructured data, due to its low change rate, is frequently the culprit in the classic, sometimes-extreme problem of not getting full backups done within the available backup window. Last year I met a potential client whose Friday night full backup job – which was mostly comprised of unstructured data – was often not finished by the time office staff started showing up for work Monday morning. Why is the low change-rate nature of unstructured data a problem as it relates to backup? Simple – every time a full backup happens, it backs up mostly the same data it backed up in the last full backup. It is an incredible waste of time, and of expensive backup storage media.
If structured and unstructured data can be treated as separate entities within the small business environment, new opportunities arise which allow for effective protection of both. I mentioned that the topic of the article is how not to backup data in the traditional sense, but I’ll qualify this slightly. I’m referring primarily to unstructured data. If unstructured data can be removed from the everyday backup schema, this relieves pressure on the systems that are charged with backing up the remaining structured data, improving the quality and speed of both backup and recovery.
How can unstructured data be adequately protected without traditional backup? There are several solutions on the Southern Data Storage line card which enable this but I’d like to highlight two in particular:
- The people at Nasuni sometimes appropriately refer to their solution as a “Data Continuity Service.” It is a filer using the CIFS or NFS protocols which can be a VM or a physical device. To the network, the Nasuni filer looks and acts just like a regular file server. What it actually does is act as a cache front-end for Amazon’s S3 cloud storage service. Data is encrypted on the filer before transport to S3, and can be easily shared among multiple remote sites, creating a highly secure “pooled” storage resource. The key thing for the purposes of this article is that the Nasuni service features a Service Level Agreement (or SLA) of 100% availability, accessibility, security and immutability. (See http://www.nasuni.com/data_services/sla for details.) This basically means that unstructured data, once stored using the Nasuni service, can be removed from the backup schema.
- Nexsan Assureon, while usually referred to as a data archiving solution, also offers a level of availability and immutability that is beyond that of typical backup systems. Assureon stores file data in a locked-down, searchable archive that can serve many purposes. It is most often used to meet requirements relating to regulatory compliance or stringent corporate data retention, but can also be used to relieve pressure on both backup and primary storage systems. Assureon deploys as a CIFS / NFS filer but is a physical storage device which can be deployed at local and replicated locations for disaster recovery purposes. (For more info, go to http://www.nexsan.com/products/assureon/.)
As I said in Part 1, I once heard someone say “Everyone has a backup application they hate – it’s usually the one they are using.” I hope I’ve been able to show that the best ways to alleviate the pain of backups are to use solutions which efficiently replicate data (see Part 1) or, for at least some data, eliminate the need for backups altogether. Hope this helps!
