IceCube Project Data Storage Requirements,2008 to 2013 IceCube 項目數(shù)據(jù)存儲要求2008 年至 2013 年_第1頁
IceCube Project Data Storage Requirements,2008 to 2013 IceCube 項目數(shù)據(jù)存儲要求2008 年至 2013 年_第2頁
IceCube Project Data Storage Requirements,2008 to 2013 IceCube 項目數(shù)據(jù)存儲要求2008 年至 2013 年_第3頁
IceCube Project Data Storage Requirements,2008 to 2013 IceCube 項目數(shù)據(jù)存儲要求2008 年至 2013 年_第4頁
IceCube Project Data Storage Requirements,2008 to 2013 IceCube 項目數(shù)據(jù)存儲要求2008 年至 2013 年_第5頁
已閱讀5頁,還剩2頁未讀, 繼續(xù)免費閱讀

下載本文檔

版權(quán)說明:本文檔由用戶提供并上傳,收益歸屬內(nèi)容提供方,若內(nèi)容存在侵權(quán),請進行舉報或認(rèn)領(lǐng)

文檔簡介

IceCubeProjectDataStorageRequirements,2008to2013

Introduction

SincetheoriginaldesignestimatesfortheIceCubeproject,whicharecapturedinthe“PreliminaryDesignDocument”,therequirementsfordatastoragehaveevolveddramatically.Forinstance,thepresentvolumeofsimulationdata,tosimulateIC22andpriorconfigurations,heldondiskislargerthanthetotalestimatefor15yearsofexperimentalandsimulationdatainthePDD.

Thisincreaseindatavolumescanbetracedtonumeroussources,howeveritisdifficulttoquantifytheeffectofeachintermsoftheircontributiontotheresultingtotalincrease.MajorincreasescanbeattributedtotheDAQ,withtheaverageeventsizebeingsignificantlyhigherthanoriginallyestimated.Thisisthenimmediatelycompoundedbythechangeinemphasistowardsalowtriggerthreshold.

Itwasalsoexpectedthatthefilteringsystemwouldbemoreadvancedthanwhatitispresently.Whilethefilteringsystemiscertainlyreachingahighlevelofmaturity,itisstillatthepointwhereithasbeennecessarytore-filterlargevolumes(manytensofTB)ofrawdata,andlargevolumesofunfilteredsimulationdataisbeingheldonline.Itmayhavebeenthattheexpectationsofbeingabletoputthissoftwarequicklyintoproductionweretoohigh.

During2007itbecameobviousthatthestoragesystemwouldnotbeabletomeettheongoingneedsoftheexperiment.Manymeasureswereattemptedtominimizetheimpactofgrowingrequirements,suchasrunningatveryhighutilizationrates,andstretchinginfrastructuretoit’slimits.Thishaspredictablyresultedinincreasedfragilityandassociateddowntime.Ithasalsoresultedinthelossoftheabilitytorespondquicklytoexpansionneedswithnotonlythediskbeingfullyutilized,butinfrastructuresuchasserversandSANswitchesatcapacity.

InresponsetothissituationameetingwasheldinNovember2007,andthenamoredetailedfollowupinJanuary2008.FromtheJanuarymeetingrequirementsforeachpresentlyidentifiedareawerecaptured,andanewanalysisstorageareaproposed.ThisdocumentwillgiveanoverviewoftheIceCubestoragetypesandareas,andpresenttherequirementsascapturedattheJanuarymeeting.

Overview

ThetotaldatastoragerequirementsforIceCubearelargebutnotmassive.Originally,whentherequirementsweremoremodest,itwasexpectedthatitwouldbepossibletostorealldataononlinedisk.Howeverithasreachedthepointthatthisisnotpracticalwithinbudgetconsiderations.AlsolargeonlinediskvolumesresultincorrespondinglyhighDataCenterspacerequirements,electrical,andcoolingneeds.Inresponsetothisitwasdecidedtointroduceatapebasedfilesystemforlongtermstorageofpotentiallylargevolumesofdata.ThissystemwasoriginallydesignedtohavelimitedfunctionalityprimarilyaimedatrestoringSouthPoledatatapes.Itisproposedthissystembeexpandedtomeetotherstorageneeds.

PresentlytheIceCubeDataCenterhasapproximately250TB(usable)ofonlinedisk,mostlyonanonspecifichardwarevendorsoftwarebasedSAN(Ibrix),andtoalesserextentdirectattachedstorage.RecentlyanHSMtapebasedfilesystemwasaddedwithaninitiallicensedcapacityof80TB.Inadditiontouseraccessiblestoragethereisabackupsystemwhichdoesnightlyincrementalbackups,andperiodicoffsitebackups.ThebackupsystemusesAtempoTimeNavigatorsoftware,andaSpectraLogicT950tapelibrary.

Theonlinestorageispresentlydividedinto3areas,experimentaldata,simulationdata,anduserdata.Ithasbeenproposedthatanadditionalstorageareaforanalysisbeadded.Thepresentstatusofonlinestorageisasfollows.

/data/exp141TB,93%used

/data/sim69TB89%used

/net/user(total)33TBbrokendownto24TB(UW)13TBused,9TB(non-UW)7.1TBused

Theuserstorageareaisarelativelysmallstorageareawhicheverycollaboratorhasaccessto.Eachuseriscurrentlylimitedto100GB,unlesstheirinstitutionprovidesfundingtoincreasethislimit,suchasthecasewithUW.ItisapracticalnecessitytohaveastorageareacloselycoupledtotheDataWarehouseforeachactiveuser,whichisessentiallyanextensionoftheirhomedirectory,whilebeingphysicallyseparatetoavoidanyperformanceissuesassociatedwithdatastorage,whichcouldimpactroutineoperationsdependantonhomedirectories.

TheIceCubestoragetypes,andassociatedresponsiblepeople,areasfollows.Thisincludestheproposedareaofanalysisstorage.

PoleDAQ:KaelHanson

OnlineFiltering:ErikBlaufuss,chairofTFTBoard

OfflineFiltering:MartinMerck

Simulation:PaoloDesiati

Analysis:GaryHill,AnalysisCoordinator

TherequirementspresentedinthisdocumenthavemostlyoriginatedatthemeetingonJanuary28th2008inMadison,whichproducedthedocument“NotesoftheStorageRequirementsMeeting”.ThisdocumentisaccompaniedbyadetailedIceCubeDataStorageRequirementsspreadsheet.

Requirements

SouthPoleDAQandOnlineFiltering

KaelHansonandErikBlaufuss

Itisproposedthatforthecomingyear,IC40,thatthecurrentarchivingarrangementcontinue,with3keytypesofDAQdataoutputbemaintained.Theseare:

Rawdatastream

DataTransferredoverthesatellite,predominatelyfiltereddata

Filtereddatanottransferredoverthesatellite

ForIC40theexpectedsizeofthesedatastreamsare,raw500GBperday(610GBuncompressed),satellite30GBperday,andfilteredbutnotoversatellite15GBperday.

ItishopedthatduringtheIC40periodthatonlinefilteringwillreachalevelofmaturitythatarchivingoftherawdatacouldbediscontinued.Infutureyearsitisproposedthattheentirerawdatastreamnotbearchived,andthattherawdataistapedforalimitedperiodaftertheadditionofnewstringsduringasettlinginperiodofthenewfilter,lessthan2months.

TheexpecteddatavolumesforIC60are750GBperday,andfiltereddata80GBperday.ForIC80theexpecteddatavolumesare1TBperdayrawdata,and135GBperdayoffiltereddata.Thedatavolumeoffiltereddatatobetransferredoverthesatellite,andfortapingforlaterphysicaltransport,willdependonupgradestotheTDRStransfersystemandconflictswithotherSouthPoleusers.In2008/2009itisexpectedthatoperationswillmovetousingTDRSF3,allowingIceCubetotransfer60GBperdayresultingin20GBperdayoffiltereddatabeingtaped.

Alldatavolumesarecompresseddata.

Theaccompanyingspreadsheetshowsbothscenariosoftaping,andnottaping,rawdata.Italsoassumesatapetechnologyupgradein2010.MoredetailsaboutSouthPoletapingiscontainedinadocument“ProposalforArchivingofDataatSouthPole”,February2008.

SimulationData

PaoloDesiati

Thedatastoragerequirementsofsimulationdependonnumerousfactorswhicharesummarizedintheattachedspreadsheet.Itwasdeterminedthattheimmediateneedsofsimulationwasanadditional10TBtocompleteIC40simulation,andmidyearanadditional50TB,andanadditional30TBinthefalltodoIC60simulation.

NextyearforIC80simulationitisexpectedthat60TBofdatacouldbemovedtoanotherstoragemediasuchastheHSM,andthatanadditional60TBwouldberequired.Thiswouldbethesteadystateforan80stringdetector.

ExperimentalData

MartinMerck

ExperimentaldataisdefinedasanydatatransferredfromtheSouthPole,andtheoutputof“PrimaryDataProcessing”(a.k.a.OfflineFiltering).Thedistinctionbetween“PrimaryDataProcessing”andAnalysisDataisbasedontheprojectswelldefinedprocessinplaceforthetwoareasmorethananythingelse.

Thedatalifecycleoffiltereddataintheofflinefilteringsystemisthatdataundergoes3processesfromtheoriginaloutputoftheonlinefiltering.TheoutputoftheonlinefilteringsystemissaidtobePFFilt.TheofflinefilteringisdoneinthenorthernhemisphereandtheoutputofeachfilteringprocessisLevel0,Level1,andLevel2.ThesizeoftheLevel2dataisapproximately50%largerthanPFFilt,andintermediatelevelsaresomewhereinbetween.

DuringtheIC40processingitisexpectedthattheoriginalPFFiltfiles,andall3outputlevelswillbeheldondiscsimultaneously.With30GBperdayofPFFiltdataarrivingoverthesatellite,thetotalrequirementofofflinefilteringwillbeapproximately105GBperday.ThustosupportofflinefilteringforIC40approximately50TBofstoragewillberequired,whichincludesasmalloverheadmargin.

AtthestartoftheIC60yeartheIC40non-satellitefiltereddatawillarrivefromSouthPoleontapes.ThisdatawillneedprocessingconcurrentlyastheIC60offlinefilteringstarts.Atthispointtheoutputofatleast2oftheIC40processinglevelscouldbedeletedprovidingthespacefortheprocessingoftheextraIC40data.TheHSMstoragesystemwouldalsobeutilizedinrestoringandstoringthenewSouthPoledata.HoweveratthisstageadditionalonlinestorageneedstobeaddedfortheoutputofIC60processing.Itisanticipatedthatatthisstagethefilteringwillbematureenoughthattheoutputofalltheprocessinglevelswillnothavetobesavedsimultaneouslyandthatanadditional60TBofstoragewouldbeadequate.ForIC80itisanticipatedthat100TBadditionalstoragewouldagainberequired.Atthispointsteadystateisreachedandanadditional100TBeachyearwouldberequiredunlessolderdatastartstobemovedtodifferentstorageareassuchastheHSM.

AnalysisData

MartinMerckandGaryHill

Itisproposedthatanewstorageareabeaddedfororganizedanalysiswithinthecollaboration.Therelevantanalysisareaswouldbeanalysisbeyondlevel2inofflinefiltering,androughlybasedontheanalysisworkinggroups.AssuchitisproposedthattheAnalysisCoordinatorwouldberesponsibleforoversightoftheusageofthisstorage,andforfuturedefinitionofrequirementsinthisarea.

Forplanningpurposesitisproposedthatthecapacityofthisareaexpandatthesamerateasofflinefiltering.Thusitstartwithaninitial50TBthisyear,expand60TBin2009,and100TBeachyearthereafter.

Accesstothisstoragewouldbebasedontheinfrastructurealreadyinplaceforsimulationandexperimentaldata.Forinstance,aleadforananalysisworkinggroupwouldcreateadirectorywithinthisstorageareaviatheDataWarehousewebinterface.Thiswouldcaptureinformationsuchasthepurposefortherequiredstorage,andaresponsibleperson.Itwouldalsoallowforsubsequenteasyaccessviathewebandquerytools.RegularusagesummarieswouldbesenttotheAnalysisCoordinator(anddesignees)foroversightofutilization.

UserData

ThereisasignificantadvantageforallusersofthecollaborationtohavingstoragecloselyconnectedtothecoreDataWarehousestorageandCPUresources.Themainpurposeofthisareaisasatemporarystorageareafordatatransferofforindividualanalysis,andhasthedesiredtechnicalconsequenceofkeepingdataoutofaccounthomedirectories,whichcancausesignificantresourceissues.

Basedonthepresentusagepatternsthecapacityseemtobeonthesmallside.Presentlyeachuserhasadefaultsoftquotaof100GB.Howevercompellingargumentshavebeenmadebynumeroususerswhichhasresultedintemporaryincreasesofupto200GB.Thecapacitywouldbeexpandedtoabitover200TBforallusers,allowinganincreaseindefaultquotasto150GBsoft,and200GBhard.Allinstitutionswouldhavetheopportunitytoaddadditionalstoragebeyondthebase,asUWhas,byanadditionalcontributiontothecommonfund.Otherwiseitisuptoindividualstoworkwithintheallocatedstorage.

Todealwithorganizedanalysiswithinthecollaborationanewstoragetype,analysisstorage,isproposed.

HSM(TapeBasedFilesystem)

AnHSMisatapedbasedfile-system.Dataisstoredontapes,butafrontenddiskbufferandapplicationmakethesystemappearasonlinedisk.Ifanattemptismadetoaccessdatanotonthebufferdisktheapplicationautomaticallyloadsthedatafromtape.Thusitisonlinestoragewithahighlatency.TheprimarypurposeoftheIceCubeHSMsystemisforaccessingthedatastoredattheSouthPole,withtheintroductionofanHSMandtapelibrarythisseasonatPole.HoweveritpresentlyhaslimitedabilitytostoredataattheDataCenterthatisnotassociatedwithSouthPole.Atpresentthemainlimitationisthatthesystemisnotsimultaneouslyreadwrite.

ItisproposedtoupgradetheHSMsystemattheDataCentertoprovideacheaperstoragetechnologyforinfrequentlyaccesseddata,aswellasprovidingaccesstotheSouthPoledata.

Archive

Archivingisatypeofstoragewhichisavailabletotheproject.Howeveritisimportanttounderstandwhatarchivingis.InthecontextofthepresentIceCubeinfrastructure(andnochangetothisisplanned)archivingisfordatathatcanpotentiallybedeletedfromaccessiblestorage,butthereissomesmallassociatedriskindoingso.Thusanarchivecopyismadejustincasethedataisneededinthefutureforunforeseenreasons.Itisnotfortemporaryremovalofdataforrestorationinthefuture.Thustherestorationofarchiveddataisnotplannedorbudgetedandanyfutureaccesswouldneedtofundedwhenrequired.Thereisalsotheissueofdatamigrationastapestoragetechnologiesadvance.Atpres

溫馨提示

  • 1. 本站所有資源如無特殊說明,都需要本地電腦安裝OFFICE2007和PDF閱讀器。圖紙軟件為CAD,CAXA,PROE,UG,SolidWorks等.壓縮文件請下載最新的WinRAR軟件解壓。
  • 2. 本站的文檔不包含任何第三方提供的附件圖紙等,如果需要附件,請聯(lián)系上傳者。文件的所有權(quán)益歸上傳用戶所有。
  • 3. 本站RAR壓縮包中若帶圖紙,網(wǎng)頁內(nèi)容里面會有圖紙預(yù)覽,若沒有圖紙預(yù)覽就沒有圖紙。
  • 4. 未經(jīng)權(quán)益所有人同意不得將文件中的內(nèi)容挪作商業(yè)或盈利用途。
  • 5. 人人文庫網(wǎng)僅提供信息存儲空間,僅對用戶上傳內(nèi)容的表現(xiàn)方式做保護處理,對用戶上傳分享的文檔內(nèi)容本身不做任何修改或編輯,并不能對任何下載內(nèi)容負(fù)責(zé)。
  • 6. 下載文件中如有侵權(quán)或不適當(dāng)內(nèi)容,請與我們聯(lián)系,我們立即糾正。
  • 7. 本站不保證下載資源的準(zhǔn)確性、安全性和完整性, 同時也不承擔(dān)用戶因使用這些下載資源對自己和他人造成任何形式的傷害或損失。

最新文檔

評論

0/150

提交評論