Forum Replies Created
-
Tim,
While everyone wants to get their archiving done as quickly as possible, I would take issue that you have to feed data to tape drives at the data rates you advise. Certainly shoeshining not only slows down writing further as you describe, but also wastes tape – I’d agree it is to be avoided.
I don’t know about IBM drives, but from what we’d been advised by HP, their drives and Quantum drives support variable write speeds and do not start shoeshining until data rates fall off severely.
Quoted values of data rate matching speeds are 46.7–140MB/s for LTO-5 and 54-160 MB/s for LTO-6.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Clones will be slightly more space efficient than disk images. Mac Disk Images (aka .dmg files) have the advantage that they are a single object that represents a disk volume, but there is some overhead in creating them. Additionally, they have to be smaller than the volume on which you save them and you have to create them at some defined size which you may not completely fill up.
Just remember that hard disk drives stored on the shelf and not regularly spun up, will eventually not be able to spin up. Normally 6 months or a year isn’t a problem, but we have heard of situations where it has been.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Hi John,
It is in fact the job of the archiving software manufacturers to implement AXF if they are going to conform to this standard.
To the best of my knowledge, the only such software manufacturer who has an end-to-end AXF solution is Front Porch Digital, and was one of the prime movers behind getting the standard adopted by SMPTE. If any of you other software manufacturers reading this know otherwise, please so advise.
IMO, at this time, AXF is a solution looking for a problem. The industry as a whole appears to be migrating from proprietary solutions to LTFS which so far is providing sufficient interchange capabilities to be the basis for most of the emerging content delivery standards.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Hi Vincent,
If you are using the Cache-A Discovery format, the metadata.xml file will be on the index partition and you really don’t need to check.
If you insist on double checking, be sure to eject the tape and reinsert after the last write to ensure that the information about where that file is written has been updated.
We recommend never deleting anything from a tape – in fact data will still be wherever it was, just no longer visible through LTFS.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Well Tim, we can agree at least on one thing, that you are long winded 😉
Actually, I don’t take issue with any of the points you make in your post above.
I was in particular making the same point as you:
[Tim Jones] “If the original data is bad on the disk, it will be bad on the tape”And then I was taking that point further – We saw many users doing extensive verification of copies of source data or even copies of copies of source data, and that is IMO a waste of time as the disk-to-disk copies are more likely to have errors than disk-to-tape.
I don’t doubt that BRU checksums are better than SHA1 or MD-5, just that they are bigger and slower. I note that no one else seems to think it is worth the overhead. And by no one else, I include IBM, HP, Quantum, Oracle, Spectra, and so on.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Every LTFS implementation (except the open-sourced free versions) support tape spanning.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Tim,
I stand by my statement that the only way to know for sure is a full restore and compare. I did not mean to imply that has to be done manually, automated means such as your Autoscan certainly meets that criteria. And I don’t dismiss checksum comparison as a powerful way to verify without a full restore – only that it should be done against the source data set. Copied data IS more likely to have bit errors going between hard drives than going onto tape.
I was being flippant, but not about user’s concern for their data. I was making my comments about the need for the level of checksumming in BRU, the significant overhead required for that, and the fact that everyone else uses much more efficient checksums such as the popular MD-5. Of course you can’t use MD-5s to reconstruct data the way BRU can, but my point is that the hardware has become so reliable you no longer need to. In saying you’ve recovered data made on other products that didn’t use verification, you have in fact reinforced that point.
If the overhead for BRU checksums is 18%, that’s not just overhead in tape cost, it’s overhead in write times, bit level verify times, and overhead in just plain having to swap, handle and store more tapes.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Tim,
I tried to provide some good general information to the forum about LTO. I’m dismayed that you focus in on one statement I made in order to reinforce your parochial outdated value proposition.
Yes, BRU has the biggest checksums in use on tape systems today… but no one ever mentions the overhead for that – LTFS can get 2.4TB on an LTO-6; how much can BRU get? How come no one else uses such large checksums?
The checksumming system BRU uses was developed to solve the problems with early tape systems which were unreliable and even with read-after-write and hardware ECC, still did indeed have unrecoverable errors. Issues like tape contact, edge pack, significant media inconsistency, and a range of mechanical flaws in early tape solutions (even as recently as LTO-1 and 2) were all good reasons to need that.
But LTO is now really reliable enough that you don’t need to verify tapes – LTO errors really are about 100 times less likely than what you get from an enterprise HDD. As your paper notes, garbage in, garbage out, and the only real checksum to compare to would be one generated on the source media, not running checksums on source data residing on any hard disk but the one that was in the camera. How many users run verification on data copied from one HDD to another (again, 100x more likely to have a bit error than copying to LTO tape)?
In my discussions with users on film sets filling multiple LTO tapes every day, they still run checksum verifications religiously but most will admit that if the tape was made without errors, they never see checksum discrepancies.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
There are many reasons why space remaining on LTO tapes may not add up the way you think:
The reason for the biggest amount of lost space I’ve seen is a result of the fact that LTO drives always do read-verified-writes at the hardware level. This means that, if a block of data did not get written properly, the tape drive will mark it as bad and rewrite the data. As tapes get old and as heads get worn or dirty, this happens more and more to the point where marked bad blocks can take up more room than the data itself. This will even happen to some extent on new drives and new tapes (more with some brands than others).
LTO tape drives all come with a built-in hardware lossless compression engine and will save an unpredictable amount of space depending upon the data. Since your case shows lost space, this is not likely to have been a big factor but will always impact to some extent the exact amount of tape space taken by any data set.
LTFS partitions each tape to keep a space for its file system (where it puts the index files). All LTO tapes are written in a serpentine manner, and LTO-5 tapes have a total of 80 tracks — 40 going from the beginning to the end and another 40 going back from the end to the beginning. The index partition is one track down plus one track back, so it is 2/80ths of the tape’s capacity or 1500GB/40 = 37.5 GB. The partition also uses another track down and back as a “guard band,” leaving 1500 – (37.5 *2) = 1425 GB useable space remaining. With LTO-6, it is 136 tracks with 2500GB capacity which works out to about the same overhead leaving 2425 GB capacity.
The size of the index will vary with the number of files but it will always be relatively small. The index is written to both the index partition and the data partition for redundancy so will take up some space. Note that each time you add to an LTFS volume, it creates a new index file and the old index files also all still reside on tape, thus many write sessions will use more space for these indices than a tape filled in one session.
Also, be aware that Macs do not represent data using the same numbers as Windows or Linux (and LTFS). A gigabyte, or GB, is now defined as 1,000 bytes cubed, or 1,000,000,000 bytes. A gibibyte, or GiB, is equal to 1024 bytes cubed, or 1,073,741,824 bytes – Mac OSX is using the decimal representation while the rest of the computing world is still using the base2 version.
LTO is very robust – if you archived without errors, the chances are extremely high (one in 10^17th) that your data is all there. The only way to know for sure however is to restore every file and compare checksums.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tomgoldberg@gmail.com
https://tomgoldberg.net -
Adam Hall is exactly right as to the reasons why this happens and if you haven’t encountered the problems it is no doubt due to the luck of the draw in how you organized the files and how you’ve traversed the file system.
Note that even if you’ve archived with baseline LTFS from a terminal, any good LTFS solution should be able to parse those tapes’ index tracks to provide a better way to look at what’s on tape as well as be able to presort restores by tape order for efficient recoveries. I know Cache-A systems do this, believe that PreRollPost should also do the same, and would hope YoYatta does as well.
Tom Goldberg
TGCS
30201 Rainbow Hill Rd.
Evergreen, CO 80439
mailto:tom@tomgoldberg.com
https://cache-a.com