




版權(quán)說(shuō)明:本文檔由用戶提供并上傳,收益歸屬內(nèi)容提供方,若內(nèi)容存在侵權(quán),請(qǐng)進(jìn)行舉報(bào)或認(rèn)領(lǐng)
文檔簡(jiǎn)介
IceCubeProjectDataStorageRequirements,2008to2013
Introduction
SincetheoriginaldesignestimatesfortheIceCubeproject,whicharecapturedinthe“PreliminaryDesignDocument”,therequirementsfordatastoragehaveevolveddramatically.Forinstance,thepresentvolumeofsimulationdata,tosimulateIC22andpriorconfigurations,heldondiskislargerthanthetotalestimatefor15yearsofexperimentalandsimulationdatainthePDD.
Thisincreaseindatavolumescanbetracedtonumeroussources,howeveritisdifficulttoquantifytheeffectofeachintermsoftheircontributiontotheresultingtotalincrease.MajorincreasescanbeattributedtotheDAQ,withtheaverageeventsizebeingsignificantlyhigherthanoriginallyestimated.Thisisthenimmediatelycompoundedbythechangeinemphasistowardsalowtriggerthreshold.
Itwasalsoexpectedthatthefilteringsystemwouldbemoreadvancedthanwhatitispresently.Whilethefilteringsystemiscertainlyreachingahighlevelofmaturity,itisstillatthepointwhereithasbeennecessarytore-filterlargevolumes(manytensofTB)ofrawdata,andlargevolumesofunfilteredsimulationdataisbeingheldonline.Itmayhavebeenthattheexpectationsofbeingabletoputthissoftwarequicklyintoproductionweretoohigh.
During2007itbecameobviousthatthestoragesystemwouldnotbeabletomeettheongoingneedsoftheexperiment.Manymeasureswereattemptedtominimizetheimpactofgrowingrequirements,suchasrunningatveryhighutilizationrates,andstretchinginfrastructuretoit’slimits.Thishaspredictablyresultedinincreasedfragilityandassociateddowntime.Ithasalsoresultedinthelossoftheabilitytorespondquicklytoexpansionneedswithnotonlythediskbeingfullyutilized,butinfrastructuresuchasserversandSANswitchesatcapacity.
InresponsetothissituationameetingwasheldinNovember2007,andthenamoredetailedfollowupinJanuary2008.FromtheJanuarymeetingrequirementsforeachpresentlyidentifiedareawerecaptured,andanewanalysisstorageareaproposed.ThisdocumentwillgiveanoverviewoftheIceCubestoragetypesandareas,andpresenttherequirementsascapturedattheJanuarymeeting.
Overview
ThetotaldatastoragerequirementsforIceCubearelargebutnotmassive.Originally,whentherequirementsweremoremodest,itwasexpectedthatitwouldbepossibletostorealldataononlinedisk.Howeverithasreachedthepointthatthisisnotpracticalwithinbudgetconsiderations.AlsolargeonlinediskvolumesresultincorrespondinglyhighDataCenterspacerequirements,electrical,andcoolingneeds.Inresponsetothisitwasdecidedtointroduceatapebasedfilesystemforlongtermstorageofpotentiallylargevolumesofdata.ThissystemwasoriginallydesignedtohavelimitedfunctionalityprimarilyaimedatrestoringSouthPoledatatapes.Itisproposedthissystembeexpandedtomeetotherstorageneeds.
PresentlytheIceCubeDataCenterhasapproximately250TB(usable)ofonlinedisk,mostlyonanonspecifichardwarevendorsoftwarebasedSAN(Ibrix),andtoalesserextentdirectattachedstorage.RecentlyanHSMtapebasedfilesystemwasaddedwithaninitiallicensedcapacityof80TB.Inadditiontouseraccessiblestoragethereisabackupsystemwhichdoesnightlyincrementalbackups,andperiodicoffsitebackups.ThebackupsystemusesAtempoTimeNavigatorsoftware,andaSpectraLogicT950tapelibrary.
Theonlinestorageispresentlydividedinto3areas,experimentaldata,simulationdata,anduserdata.Ithasbeenproposedthatanadditionalstorageareaforanalysisbeadded.Thepresentstatusofonlinestorageisasfollows.
/data/exp141TB,93%used
/data/sim69TB89%used
/net/user(total)33TBbrokendownto24TB(UW)13TBused,9TB(non-UW)7.1TBused
Theuserstorageareaisarelativelysmallstorageareawhicheverycollaboratorhasaccessto.Eachuseriscurrentlylimitedto100GB,unlesstheirinstitutionprovidesfundingtoincreasethislimit,suchasthecasewithUW.ItisapracticalnecessitytohaveastorageareacloselycoupledtotheDataWarehouseforeachactiveuser,whichisessentiallyanextensionoftheirhomedirectory,whilebeingphysicallyseparatetoavoidanyperformanceissuesassociatedwithdatastorage,whichcouldimpactroutineoperationsdependantonhomedirectories.
TheIceCubestoragetypes,andassociatedresponsiblepeople,areasfollows.Thisincludestheproposedareaofanalysisstorage.
PoleDAQ:KaelHanson
OnlineFiltering:ErikBlaufuss,chairofTFTBoard
OfflineFiltering:MartinMerck
Simulation:PaoloDesiati
Analysis:GaryHill,AnalysisCoordinator
TherequirementspresentedinthisdocumenthavemostlyoriginatedatthemeetingonJanuary28th2008inMadison,whichproducedthedocument“NotesoftheStorageRequirementsMeeting”.ThisdocumentisaccompaniedbyadetailedIceCubeDataStorageRequirementsspreadsheet.
Requirements
SouthPoleDAQandOnlineFiltering
KaelHansonandErikBlaufuss
Itisproposedthatforthecomingyear,IC40,thatthecurrentarchivingarrangementcontinue,with3keytypesofDAQdataoutputbemaintained.Theseare:
Rawdatastream
DataTransferredoverthesatellite,predominatelyfiltereddata
Filtereddatanottransferredoverthesatellite
ForIC40theexpectedsizeofthesedatastreamsare,raw500GBperday(610GBuncompressed),satellite30GBperday,andfilteredbutnotoversatellite15GBperday.
ItishopedthatduringtheIC40periodthatonlinefilteringwillreachalevelofmaturitythatarchivingoftherawdatacouldbediscontinued.Infutureyearsitisproposedthattheentirerawdatastreamnotbearchived,andthattherawdataistapedforalimitedperiodaftertheadditionofnewstringsduringasettlinginperiodofthenewfilter,lessthan2months.
TheexpecteddatavolumesforIC60are750GBperday,andfiltereddata80GBperday.ForIC80theexpecteddatavolumesare1TBperdayrawdata,and135GBperdayoffiltereddata.Thedatavolumeoffiltereddatatobetransferredoverthesatellite,andfortapingforlaterphysicaltransport,willdependonupgradestotheTDRStransfersystemandconflictswithotherSouthPoleusers.In2008/2009itisexpectedthatoperationswillmovetousingTDRSF3,allowingIceCubetotransfer60GBperdayresultingin20GBperdayoffiltereddatabeingtaped.
Alldatavolumesarecompresseddata.
Theaccompanyingspreadsheetshowsbothscenariosoftaping,andnottaping,rawdata.Italsoassumesatapetechnologyupgradein2010.MoredetailsaboutSouthPoletapingiscontainedinadocument“ProposalforArchivingofDataatSouthPole”,February2008.
SimulationData
PaoloDesiati
Thedatastoragerequirementsofsimulationdependonnumerousfactorswhicharesummarizedintheattachedspreadsheet.Itwasdeterminedthattheimmediateneedsofsimulationwasanadditional10TBtocompleteIC40simulation,andmidyearanadditional50TB,andanadditional30TBinthefalltodoIC60simulation.
NextyearforIC80simulationitisexpectedthat60TBofdatacouldbemovedtoanotherstoragemediasuchastheHSM,andthatanadditional60TBwouldberequired.Thiswouldbethesteadystateforan80stringdetector.
ExperimentalData
MartinMerck
ExperimentaldataisdefinedasanydatatransferredfromtheSouthPole,andtheoutputof“PrimaryDataProcessing”(a.k.a.OfflineFiltering).Thedistinctionbetween“PrimaryDataProcessing”andAnalysisDataisbasedontheprojectswelldefinedprocessinplaceforthetwoareasmorethananythingelse.
Thedatalifecycleoffiltereddataintheofflinefilteringsystemisthatdataundergoes3processesfromtheoriginaloutputoftheonlinefiltering.TheoutputoftheonlinefilteringsystemissaidtobePFFilt.TheofflinefilteringisdoneinthenorthernhemisphereandtheoutputofeachfilteringprocessisLevel0,Level1,andLevel2.ThesizeoftheLevel2dataisapproximately50%largerthanPFFilt,andintermediatelevelsaresomewhereinbetween.
DuringtheIC40processingitisexpectedthattheoriginalPFFiltfiles,andall3outputlevelswillbeheldondiscsimultaneously.With30GBperdayofPFFiltdataarrivingoverthesatellite,thetotalrequirementofofflinefilteringwillbeapproximately105GBperday.ThustosupportofflinefilteringforIC40approximately50TBofstoragewillberequired,whichincludesasmalloverheadmargin.
AtthestartoftheIC60yeartheIC40non-satellitefiltereddatawillarrivefromSouthPoleontapes.ThisdatawillneedprocessingconcurrentlyastheIC60offlinefilteringstarts.Atthispointtheoutputofatleast2oftheIC40processinglevelscouldbedeletedprovidingthespacefortheprocessingoftheextraIC40data.TheHSMstoragesystemwouldalsobeutilizedinrestoringandstoringthenewSouthPoledata.HoweveratthisstageadditionalonlinestorageneedstobeaddedfortheoutputofIC60processing.Itisanticipatedthatatthisstagethefilteringwillbematureenoughthattheoutputofalltheprocessinglevelswillnothavetobesavedsimultaneouslyandthatanadditional60TBofstoragewouldbeadequate.ForIC80itisanticipatedthat100TBadditionalstoragewouldagainberequired.Atthispointsteadystateisreachedandanadditional100TBeachyearwouldberequiredunlessolderdatastartstobemovedtodifferentstorageareassuchastheHSM.
AnalysisData
MartinMerckandGaryHill
Itisproposedthatanewstorageareabeaddedfororganizedanalysiswithinthecollaboration.Therelevantanalysisareaswouldbeanalysisbeyondlevel2inofflinefiltering,androughlybasedontheanalysisworkinggroups.AssuchitisproposedthattheAnalysisCoordinatorwouldberesponsibleforoversightoftheusageofthisstorage,andforfuturedefinitionofrequirementsinthisarea.
Forplanningpurposesitisproposedthatthecapacityofthisareaexpandatthesamerateasofflinefiltering.Thusitstartwithaninitial50TBthisyear,expand60TBin2009,and100TBeachyearthereafter.
Accesstothisstoragewouldbebasedontheinfrastructurealreadyinplaceforsimulationandexperimentaldata.Forinstance,aleadforananalysisworkinggroupwouldcreateadirectorywithinthisstorageareaviatheDataWarehousewebinterface.Thiswouldcaptureinformationsuchasthepurposefortherequiredstorage,andaresponsibleperson.Itwouldalsoallowforsubsequenteasyaccessviathewebandquerytools.RegularusagesummarieswouldbesenttotheAnalysisCoordinator(anddesignees)foroversightofutilization.
UserData
ThereisasignificantadvantageforallusersofthecollaborationtohavingstoragecloselyconnectedtothecoreDataWarehousestorageandCPUresources.Themainpurposeofthisareaisasatemporarystorageareafordatatransferofforindividualanalysis,andhasthedesiredtechnicalconsequenceofkeepingdataoutofaccounthomedirectories,whichcancausesignificantresourceissues.
Basedonthepresentusagepatternsthecapacityseemtobeonthesmallside.Presentlyeachuserhasadefaultsoftquotaof100GB.Howevercompellingargumentshavebeenmadebynumeroususerswhichhasresultedintemporaryincreasesofupto200GB.Thecapacitywouldbeexpandedtoabitover200TBforallusers,allowinganincreaseindefaultquotasto150GBsoft,and200GBhard.Allinstitutionswouldhavetheopportunitytoaddadditionalstoragebeyondthebase,asUWhas,byanadditionalcontributiontothecommonfund.Otherwiseitisuptoindividualstoworkwithintheallocatedstorage.
Todealwithorganizedanalysiswithinthecollaborationanewstoragetype,analysisstorage,isproposed.
HSM(TapeBasedFilesystem)
AnHSMisatapedbasedfile-system.Dataisstoredontapes,butafrontenddiskbufferandapplicationmakethesystemappearasonlinedisk.Ifanattemptismadetoaccessdatanotonthebufferdisktheapplicationautomaticallyloadsthedatafromtape.Thusitisonlinestoragewithahighlatency.TheprimarypurposeoftheIceCubeHSMsystemisforaccessingthedatastoredattheSouthPole,withtheintroductionofanHSMandtapelibrarythisseasonatPole.HoweveritpresentlyhaslimitedabilitytostoredataattheDataCenterthatisnotassociatedwithSouthPole.Atpresentthemainlimitationisthatthesystemisnotsimultaneouslyreadwrite.
ItisproposedtoupgradetheHSMsystemattheDataCentertoprovideacheaperstoragetechnologyforinfrequentlyaccesseddata,aswellasprovidingaccesstotheSouthPoledata.
Archive
Archivingisatypeofstoragewhichisavailabletotheproject.Howeveritisimportanttounderstandwhatarchivingis.InthecontextofthepresentIceCubeinfrastructure(andnochangetothisisplanned)archivingisfordatathatcanpotentiallybedeletedfromaccessiblestorage,butthereissomesmallassociatedriskindoingso.Thusanarchivecopyismadejustincasethedataisneededinthefutureforunforeseenreasons.Itisnotfortemporaryremovalofdataforrestorationinthefuture.Thustherestorationofarchiveddataisnotplannedorbudgetedandanyfutureaccesswouldneedtofundedwhenrequired.Thereisalsotheissueofdatamigrationastapestoragetechnologiesadvance.Atpres
溫馨提示
- 1. 本站所有資源如無(wú)特殊說(shuō)明,都需要本地電腦安裝OFFICE2007和PDF閱讀器。圖紙軟件為CAD,CAXA,PROE,UG,SolidWorks等.壓縮文件請(qǐng)下載最新的WinRAR軟件解壓。
- 2. 本站的文檔不包含任何第三方提供的附件圖紙等,如果需要附件,請(qǐng)聯(lián)系上傳者。文件的所有權(quán)益歸上傳用戶所有。
- 3. 本站RAR壓縮包中若帶圖紙,網(wǎng)頁(yè)內(nèi)容里面會(huì)有圖紙預(yù)覽,若沒(méi)有圖紙預(yù)覽就沒(méi)有圖紙。
- 4. 未經(jīng)權(quán)益所有人同意不得將文件中的內(nèi)容挪作商業(yè)或盈利用途。
- 5. 人人文庫(kù)網(wǎng)僅提供信息存儲(chǔ)空間,僅對(duì)用戶上傳內(nèi)容的表現(xiàn)方式做保護(hù)處理,對(duì)用戶上傳分享的文檔內(nèi)容本身不做任何修改或編輯,并不能對(duì)任何下載內(nèi)容負(fù)責(zé)。
- 6. 下載文件中如有侵權(quán)或不適當(dāng)內(nèi)容,請(qǐng)與我們聯(lián)系,我們立即糾正。
- 7. 本站不保證下載資源的準(zhǔn)確性、安全性和完整性, 同時(shí)也不承擔(dān)用戶因使用這些下載資源對(duì)自己和他人造成任何形式的傷害或損失。
最新文檔
- 規(guī)范貨物貿(mào)易規(guī)則
- Unit+5+I+think+that+mooncakes+are+delicious同步練-+2024-2025學(xué)年魯教版(五四學(xué)制)八年級(jí)英語(yǔ)下冊(cè)+
- 2025年教師招聘考試教育學(xué)心理學(xué)選擇題復(fù)習(xí)題庫(kù)
- 2024年上海市浦東新區(qū)中考二模語(yǔ)文試卷含詳解
- 2025年渭南貨運(yùn)從業(yè)資格證模擬考試
- 2025年湖南貨運(yùn)車從業(yè)考試題
- 貸款行業(yè)客戶經(jīng)理經(jīng)驗(yàn)分享
- 2025勞動(dòng)合同解除協(xié)議書(shū)范本
- 2025企業(yè)兼職財(cái)務(wù)顧問(wèn)合同協(xié)議書(shū)
- 2025辦公室租賃合同附加協(xié)議書(shū)
- 《孫權(quán)勸學(xué)》歷年中考文言文閱讀試題40篇(含答案與翻譯)(截至2024年)
- 《高速鐵路系統(tǒng)》課件
- 新型可瓷化膨脹防火涂料的制備及性能研究
- 《機(jī)械設(shè)計(jì)課程設(shè)計(jì)》課程標(biāo)準(zhǔn)
- 肺結(jié)核防治知識(shí)培訓(xùn)課件
- 《新生兒沐浴和撫觸》課件
- 《基于作業(yè)成本法的S公司物流成本分析研究》8300字(論文)
- 2024-2030年中國(guó)負(fù)載均衡器行業(yè)競(jìng)爭(zhēng)狀況及投資趨勢(shì)分析報(bào)告
- 浙江省溫州市重點(diǎn)中學(xué)2025屆高三二診模擬考試英語(yǔ)試卷含解析
- 電力工業(yè)企業(yè)檔案分類表0-5
- 臨時(shí)用地草原植被恢復(fù)治理方案
評(píng)論
0/150
提交評(píng)論