Saturday, April 24, 2010

Reducing Carbon Foot Print in IT

Carbon is fast becoming a currency that global organizations cannot ignore. With carbon cap-and-trade schemes either being planned or implemented by a growing number of national and regional authorities, the carbon impact of an organization’s operations is becoming a measurable cost. And after the global recession, with the focus still very much on the bottom line, carbon impact management is closer than ever to becoming a universal boardroom issue. But who, within an organization, is responsible for mitigating the carbon cost? Certainly, this question does not yet have a straightforward answer. But one thing is clear – however organizations choose to deal with the carbon question, CIOs are bound to be involved. This is because – even if they are not given primary responsibility for managing organizational carbon impact – CIOs will need to ensure carbon reporting systems remain up to date with legislative requirements. On top of that, there is a growing realization that ICT itself has a significant role to play in carbon impact management. The European Commission recently announced1 the information and communication technologies (ICT) sector should lead the transition to an energy efficient economy. It called for Europe’s ICT sector to:

• Agree on common energy consumption measures
• Overtake the EU’s 2020 targets by 2015
• Make innovative use of ICT to make Europe a low-carbon economy

The EC said replacing 20 per cent of European business trips by video conferencing could save more than 22 million tons of CO2 per year. It also said that broadband facilitating increased use of online public services could save two per cent of total worldwide energy use by 2020. It is clear that CIOs need at least to know all the facts, if they are to make an informed decision about the role they will play. So where should they begin?

Carbon Emission and the Bottom line:-
The poster child in the war against carbon emissions has been ‘green’ energy. But while the likes of hydropower, biomass, wind and solar energy may have a knack of exciting the headline writers, none have yet become affordable mainstream technologies. And while the race to find ways to replace our reliance on ‘dirty’ fuels needs to go on, the place to look for short-term emission cuts is in energy efficiency. Unglamorous it may be, but 40 per cent of the carbon reduction to be achieved by 2020 and beyond needs to come from precisely this source2. And central to the story of how business will meet those targets is IT. When asked to provide an example of environmental irresponsibility, the airline industry is never far from people’s lips. Yet global CO2 emissions from IT are roughly on a par with those pumped into the atmosphere by planes. What’s more, the opportunities to use technology to cut emissions are vast. A recent report notes that IT could contribute as much as 15 per cent of global emissions reductions by 20203. Of course, the idea of cutting the energy consumption associated with IT is not new. It’s a target that has featured in the CSR programs of big business for some time. What is new however is the growing realization that an energy efficient approach to IT is far more than a PR tool. Instead it is a bottom-line issue. It doesn’t just look good on the corporate website, it can help save serious sums of money. Consider the 15 per cent figure mentioned above. That equates to €600 billion in cost savings.

After Copenhagen:-
When world leaders met to discuss climate change at the Copenhagen Summit, carbon reduction was one of the topics under discussion. It played a central role in the Copenhagen Accord, which was the key output from the summit (for more details, see Reference Section “Copenhagen: implications for global business”). As part of the Accord, countries were invited to submit their own carbon reduction targets by the end of January 2010. Fifty five did so. They include the US, all EU countries and China, as well as major emerging economies such as Brazil, Indonesia and India. Between them these nations emit 78 per cent of the world’s greenhouse gases. So although there were notable absentees – Brazil was the only South American nation to volunteer a pledge, and just six out of 55 African countries did so – and although the combined country targets are insufficient to cap the temperature rise at the desired two degrees, this was still an important step on the path towards ultimately achieving a legally-binding global agreement. (Though just when such a step will be taken is another matter.) Crucially, the Accord also provides for scrutiny to monitor whether or not countries meet their emissions reductions targets – a key point of difference between developing and developed countries throughout the negotiations. Unsurprisingly this was an issue that the US was particularly keen on in relation to China. As The Economist puts it: “Unless China can be shown to live up to its promises, it will be very difficult to get a climate bill through America’s Senate.”

Local Action – A Snapshot of Carbon reduction activity around the globe:-

Companies will be required to measure all electricity, gas, and oil use (excluding transport and travel). They will then purchase allowances equal to their annual emissions. Within that overall limit, individual organizations can decide on the most cost-effective way to reduce their emissions. Just like most cap-and-trade schemes they will then have the ability to buy extra allowances or invest in ways to cut the number of allowances they need to buy.
A league table is then created, with credits being handed out based on that year’s performance. An organization at the top of the table will receive repayments totaling more than has been paid for the allowances in the first place, while those at the bottom will receive repayments that are less than the amount paid out. In other words, organizations in the bottom half will lose money. Analysis Mason estimates that for the biggest companies, being ranked at the bottom of the table could amount to a financial penalty of over £120,0007.

How IT can cut Enterprise Cost:-
Several technologies are leading the way in helping businesses save hundreds of millions of pounds in the process of cutting their emissions.

Green data centers
One of the most exciting is data centre virtualization, a technology that slashes the number of servers required to run your organization. In the case of BT, data centres used to account for a significant chunk of the company’s carbon emissions. However the average server is utilised at a tiny percentage of its overall capacity. By ‘virtualising’ these servers, or by asking each server to carry out an increased number of tasks at the same time, utilization shoots up, and this means the number of servers needed slumps. In one of BT’s data centers, the number dropped from 1,500 to around 100, saving £600,000. Large numbers of servers create lots of heat which, in turn, means the need for power-hungry air conditioning systems. BT has redesigned its data centers to allow fresh air to cool servers. Air conditioning is only used on the rare occasions when the temperature reaches 28°C or above.

Flexible working and home working
Flexible working and home working are more familiar approaches to cutting power use(and improving productivity) yet many organizations still make little use of them. At a time when there’s an ongoing pressure to find cost savings, there’s a strong argument to take a fresh look at the efficient use and potential rationalisation of building space. BT’s own results make a persuasive case. By enabling over 13,700 employees to become home workers the company saves €750 million a year in property management costs.

Conferencing and collaboration tools
Using conference calls to replace face to face meetings has had similarly dramatic effects within BT, saving the company an estimated £183 million in 2008 and saving over 50,000 tonnes of CO2.
Collaboration technologies have the same effect, allowing people to work together without the need to travel to the same location.

Tips to Decarbonate IT:-
CIOs are at the heart of the emissions reduction story. Information technology is not just a huge energy consumer but can also play a role in reducing energy use and cutting costs in other areas of the organisation. But what are the first steps CIOs can take today to ensure they are prepared to lead the way in cutting their carbon impact, as soon as it becomes a bottom-line issue within their organization?

1. Build environmental responsibility into your distributed organization
As organizations have globalised in recent decades, they have had to face the challenge of maintaining a common corporate culture, irrespective of geography. They now face the same challenge in their attempts to engender a sense of carbon impact responsibility across the business. Concerted and consistent education, support and encouragement are needed to ensure people are aware of what is expected of them, and that they actually exhibit the desired behavior. But this message needs to come from the top down – the CIO needs to lead the way.

2. Ensure the right systems and processes are in place
A group-wide function should be put in place to capture, manage and report all the data related to IT use in the business. Work with facilities management, environmental management and finance, because they will each potentially be affected by legislation as it is brought in around the world. With more countries looking to enshrine environmental commitments in law, there is an increasing need to display accountability at the highest level of the organization. Simply storing information on a spreadsheet will no longer suffice.

3. Redefine the workplace
The vast majority of action on climate change so far has focused on energy use. Surprisingly there has been very little emphasis on travel, whether between business locations, or from home to the traditional workplace. In the UK, for example, approximately 25% of all emissions are travel-related. Promoting technologies that can replace the need for travel (while also saving huge sums of money), by enabling people to meet virtually or work nomadically, is something that CIOs can be doing at board level.

Appendix:-
Copenhagen: implications for global businesses
“It’s very disappointing, I would say, but it is not a failure.” So said Sergio Serra, Brazil’s Climate Change Ambassador at the conclusion of the Copenhagen Summit. This was a view echoed by many, from fellow politicians, to NGOs, to the legion of green pundits in the world’s media. After the summit, there was a feeling that at least a global consensus on climate change was reached. The Copenhagen Accord gives international backing for an immediate global move towards action on climate change. In some ways this is a bigger achievement than the binding commitments that resulted from Kyoto 17 years earlier when the deal only affected developed countries. So what exactly did they agree to? There were three key components:
1. Backing for an overall limit on global warming of two degrees;
2. Agreement that all countries need to take action on climate change;
3. The provision of $30 billion of immediate short-term funding from developed countries over the next three years to kick start emission reduction measures and help the poorest countries adapt to the impacts of climate change; as well as a commitment by developed countries to provide long-term
financing of $100 billion a year by 2020.

Illustration: Comparison of emissions for a typical 200-server Windows network



This illustrative comparison shows the major impact virtualization and data centre energy efficiency can make to CO2 emissions.
1. Impact of virtualization
• Non virtualised: 130KW (~610 tons CO2)
• Virtualised: 24-50KW (~112 – 130 tons CO2)
• Virtualization can reduce power consumption by between 66 per cent and 82 per cent.
(Assumption: Average data centre Power Usage Efficiency* of 2.4)
2. Impact of data centre efficiency
• The difference in energy consumption between an inefficient data centre and an efficient one can be as much as 50 per cent. (Assumption: PUE range from 3.2 to 1.6)
3. Combined impact:
Best and worst-case CO2 emission scenarios


• Potential energy saving/CO2 saving ~91 per cent
* Note: PUE is the industry standard energy efficiency measure for data centers.

References:-
1. www.smart2020.org
2. Analysys Mason’s Carbon Reduction Commitment Brochure
(http://www.analysysmason.com/PageFiles/13899/Carbon_Reduction_Commitment_Brochure.pdf)
3. www.computerweekly.com
4. EU Press Release

MINAL CHINCHKAR

Wednesday, April 21, 2010

Cloud based Fix it center for automated troubleshooting

Microsoft has just released Fix it center online beta and it is based on the automated troubleshooter feature which we saw inbuilt in windows 7. However as this is more of cloud based fix it center, it supports various operating system versions given below. Though it is a great way of managing all your PC from one location with an automated repair functionality of known issues, however if right now there is no troubleshooter for your problem, you will be provided various other related articles and community groups too. you also have an option of raising a support request from the same online portal if nothing mentioned above helped you. I think this tool is very good for individual customers and small business customers.


• Windows XP SP3

• Windows XP Pro (64-bit) SP2

• Windows Vista

• Windows 7

• Windows Server 2003 SP2

• Windows Server 2008

• Windows Server 2008 R2

You can use any computer with Internet connection to get started with Fix it Center. Simply download the Fix it Center client and follow the on-screen instructions to complete the setup. You can install Fix it Center client on as many PCs you like. It is recommend to sign up for Fix it Center Online during setup so you can manage all your computers from a single location on the Internet yet can view solutions specific for each PC. With automated troubleshooters, Fix it Center helps solve issues with your PC, even if you're not sure what the exact problem is. Fix It Center scans your device to diagnose and repair problems, then gives you the option to "Find and fix" or to "Find and report. With a single view of all your devices, it’s easy to manage multiple devices from one view. You can even manage them remotely.

I downloaded the fix it agent and ran on my PC and as you see below its running an
inventory for my PC

 
It asks me to create an online account and i can use my existing live/hotmail id.
 
 
 
 
 

I see some automated troubleshooters on the screen and the remaining are on the bottom right corner where it says " having a different problem"
 
 
I can manage multiple PC's from same console and also see the diagnostic history on all PC's. I can find more solutions from the Tab and if i need to raise a support request to microsoft, i can do it from here and can also see the old support requests.
 
 
Though i see this tool as a great way of automating troubleshooting and managing small environments, I dont think i will ever run this on my servers as of now. Hope you like this new tool as it automates troubleshooting and make life easier for an end individual user or a small business user.
 
GAURAV ANAND

Sunday, April 18, 2010

How Dynamic memory Feature of 2008 R2 Sp1 works with Failover Clustering

The news of win2k8 R2 SP1 has started coming online and the most interesting feature is dynamic memory support for Hyper-V. Constraints on the allocation of physical memory represents one of the greatest challenges organizations face as they adopt new virtualization technology and consolidate their infrastructure. With Dynamic Memory, an enhancement to Hyper-v introduced in Windows Server 2008 R2 SP1, organizations can now make the most efficient use of available physical memory, allowing them to realize the greatest possible potential from their virtualization resources. Dynamic Memory allows for memory on a host machine to be pooled and dynamically distributed to virtual machines as necessary. Memory is dynamically added or removed based on current workloads, and is done so without service interruption. At a high level, Hyper-V Dynamic Memory is a memory management enhancement for Hyper-V designed for production use that enables customers to achieve higher consolidation/VM density ratios.

Dynamic Memory is supported on guest machines running Windows Server 2003 SP2 Datacenter or Enterprise editions, Windows Server 2008 SP1 Datacenter or Enterprise editions, Windows Server 2008 R2 Datacenter or Enterprise editions, Windows Vista SP2 Enterprise or Ultimate editions, and Windows 7 Enterprise or Ultimate editions. In today’s blog I am going to focus on how dynamic memory features along with failover clustering is a win win for everyone. I am going to take example of a 3 node cluster running on 2008 R2 server core with 16 GB ram each where Red VM’s workload require 4GB ram and Grey VM’s workload require 3GB ram.


So in case of static memory assignment we see that if a cluster node goes down I either need a passive node or I need to reduce the desnsity of VM’s so that that I can place VM’s from failing node on my existing cluster nodes. In the present case [shown above] I have reduced the density of VM that I can place on Cluster nodes and by doing that I am actually loosing out usage of 5GB ram on each cluster host. However when one of the node goes down the other 2 nodes will be able to take the load of the exiting node. So lets see what happens if one of my nodes actually go down for a short interval of time.


Though I see that my virtual machine resources do failover to node A and node B, yellow VM can still not come online on node B till the time I change its static ram and reduce it so that atleast it can start up on node B. This has to be a manual operation however even by doing that I cannot guarantee that workload running on that VM will have a sufficient performance. so in ideal scenario I should modify all the VM’s ram running on node B so that my yellow VM gets enough ram to give me satisfactory performance with provided resources in the short time window where one of my node is down temporarily. But the problem here is that I cannot reduce ram on fly from other VM’s without shutting them down which I cannot as they are in production and shutting them down is not feasible. Here, Dynamic memory assignment which is a feature of 2008 R2 Sp1 brings in the magic.
 
Now we have all 2008 R2 server core Nodes with Sp1 installed and hence dynamic memory assignment feature is available/VM. For every highly available VM running on all the 3 nodes I have given 512 MB as the initial memory [customizable] and 4 GB as the max memory. Dynamic memory allows us to configure a virtual machine so that the amount of memory assigned to the virtual machine is adjusted while the virtual machine is running, in comparison to the amount of memory that is actually being used by the virtual machine. This allows us to run a higher number of virtual machines on a given physical node. It also ensures that memory is always distributed optimally between running virtual machines. Some applications assign fixed amounts of memory based on the amount of memory available when the application first starts at start of operating system. These applications will perform better with higher values for the initial memory instead of 512 mb ram that I have assigned so it depends on the workload that you are running in VM. So in our scenario, if node C now goes down, the highly available virtual machines resources will failover to other nodes and will come online with optimized host ram being shared between Virtual machines and no manual intervention required. The great benefit that comes with dynamic memory is that if a VM do not require ram at one point of time, that ram can be leverage by other VM and when required can be given back. This allows us to increase VM density, better usage of ram resources and much more control and optimization of ram resources in a failover cluster node failure scenario.
 
 
There’s lot more coming about dynamic memory here and very soon we will also compare this with VMware overcommit/ballooning feature when we get more light thrown on how it works. Also very recently XEN has released new version of their hypervisor will additional features like Transcendent Memory and Page Sharing in Xen 4.0 to enhance the performance and capabilities of the hypervisor memory operations. Xen 4.0 now supports live transactional synchronization of VM states between physical servers [very similar to VMware fault tolerance feature]. Keep checking our blog for upcoming articles and thanks for your time and hope this blog would have been a reading pleasure for you and an insight into very nice dynamic memory feature coming in 2008 R2 sp1 which is yet to be released.
 
GAURAV ANAND

Friday, April 16, 2010

Are we molding new Technology to suit ourselves or getting molded by new Technology


The phone that you see on the right side is my favorite phone and I have purchased 3 of those in last 6 years and still using same phone very happily. Once I switched to the new age smart phones but just in a month I gifted it to my mother as I figured out that usage of that new age smart phone is actually becoming a disadvantage for me. How can a cool phone will all great nice features and different apps can be a disadvantage?..ok every one is a different personality and I figured out in one month that instead of saving time this phone is making me spending lot of time in clicking waste snaps [which I would not have clicked otherwise from my semi DSLR camera], listening much more music then I usually used to do, wasting money and time in downloading apps, songs, graphics, games. I configured outlook on my mobile to be bogged down by emails even after my office hours, the thick line between my office and home time started shrinking. The apps like facebook, twitter and msn started keeping me engaged as they were accessible to me 24 hours in my hand. Nothing wrong in using social networking applications, clicking snaps, playing games, listening songs and I use all this today too but I have a dedicated nice laptop with nice gaming hardware, a nice camera to pursue my photography hobby, good surround sound speaker/home theatre for entertainment …the difference is that this allows me to have a control of my time and mind and I still can do all what I want to do without compromising on any of my experiences. Personally with that old phone I need not to care about it being stolen [ I have forgotten this phone maytimes and people have returned me or it was found where it was left ] and nor I worry about its wear and tear…in other words it an economical phone which does all that a phone should do…and on top of it …it helps me in time management and more control over my time without any time maanagement software in it. However this does not mean that you should not use those high end phones but you should know whether you really need them and utilizing them or are they just with you for your coolness quotient. It is fine if those phones are serving you and not you serving them.

Now why I wrote all this!!....
 
Many times I have seen people seeking advice on shall we upgrade to latest version?...shall we purchase latest hardware?....shall we switch to this new technology which is buzzing in market?.....cloud……Virtualization! …. Are we molding new Technology to suit ourselves or getting molded by new Technology. Do not switch over to new technology just because it worked for others in their enironment…make sure it works for you, do your own prrof of concept, bench marking and base line testings before you make any decision just based on sales presentations. There is no one technology solution that works for any IT Enterprise and that’s where identifying your buiness needs and aligning them with your existing processes and architecture come in picture. Hope you liked this analogy and it made you think for a minute and once again thanks for staying with our blog.
 
GAURAV ANAND

Tuesday, April 13, 2010

Key Home Takeways from my First day of Tech ED India 2010


I have just returned from the Hotel Lalit Ashok, Bangalore, where Tech ED 2010 event is happening this year. I am going to talk about the key home takeaways from the first day of Tech Ed 2010. It was a fabulous start with visual studio 2010 launch event however as I am not a developer I will talk more about architecture track where we discussed about different cloud patterns and practices which was followed by discussion on dynamic data center took kit session. Post lunch we saw a session on system center service manager which let enterprise admin to follow the ITIL practices and automatically generate an incident based on the alerts from SCOM. Also a self service portal can be created where a user request resources/applications and after approval from his manager those resources/applications will automatically get deployed via SCCM. There was another session on provisioning virtual machines using system center virtual machine manager via creating a hardware and software templates. SCVMM lets end user to self provision virtual machines as per his requirements. Then we jumped into another session where we saw the capability to manage linux and unix servers from SCOM. It’s a great feature which comes with SCOM 2007 R2…so you got default management packs for unix and linux flavors which enable you to generate alerts in SCOM and manage heterogeneous environments in your datacenter. One of the last sessions before demo extravaganza was for protecting virtual machines from DPM 2010 where we specifically deep dived into Cluster shared volumes [CSV]. We saw the DPM backing up virtual machine irrespective of live migration of storage migration happening on them. Best practices of using hardware vss provider for protecting CSV luns was discussed and we also saw a demo where we can recover a VM on a hyper-v host other than from which it was backed up. During breaks I also managed to visit Citrix partner demo tent where 2 Citrix employees where kind enough to show me Citrix xen desktop and Citrix xen server solution and the discussion got stretched as we compared hyper-V VDI and Citrix VDI solution from a 100 FT view. The nice thing about xen server is that the free edition contains features like live migration and Xen motion [synonymous to storage migration/V-motion]. You just need to install the xen server on a bare metal machine and then you install a client app on any other machine [synonymous to VMware Virtual center and MS SCVMM] to manage the xen server and this is again free of cost however yes the High availability features needs to be bought but still I think that’s a great technology which Citrix is providing to SMB users for free. Citrix already is delivering HDX [High definition experience] technology which Microsoft is about to roll out in 2008 R2 sp1 known as Remote FX. New HDX media stream and HDX plug and play capabilities ensure that whether users are watching windows multimedia or connecting multiple monitors, scanners, phone or other device they get same experience of local pc on a hosted virtualized desktop.

In the end the demo extravaganza was fun, where some very cool technologies and features were demoed including windows phone 7 and surface computer. We saw some cool features of power point 2010 where you can do a lot of magic by embedding your videos in power point and editing them directly from power point and same for the images. We saw how we can play with websites HTML code using IE 8 developer tools. Another cool feature was chatting/emailing in office communicator/outlook in Indian local languages. We saw some great features of search in Bing and how its beat Google in those features. It was great fun and that’s all for now. Looking forward to day 2. Thanks and good bye till I catch you tomorrow.

Gaurav Anand

Monday, April 5, 2010

All about maintennance mode, chkdsk/defrag in failover clustering including "cluster shared volumes"

Continuing from “storage architecture changes” to “how configuring storage with volume guid works”, today we will talk about different options of running chkdsk/defrag in maintenance mode in failover clustering 2008/R2. We will also see how putting a physical disk resource in maintenance mode is different from putting a cluster shared volume in maintenance mode. Chkdsk has always been a challenge especially in environments where storage disks size go in terabytes. I have seen many real life scenarios where critical disk resources were not available for services and applications in production time as chkdsk started running on them when they were mounted [brought online]…this happens because when we bring a disk resource online we check the dirty bit on file system to see if it’s dirty or not and if found dirty..we start chkdsk on it…what to do?....let’s call Microsoft….and as recommended they say “ it’s not recommended to stop chkdsk while its running”…ok fair enough..same question..what to do? ….the disk resource size is in terabytes and it may take hours if not days to get it finished. We cannot let our production critical apps/services waiting on chkdsk for days.

You have 2 options…loose production and wait for chkdsk to finish..kill chkdsk from task manager on your own risk [assuming you have all data backup]. And here comes the role of designing, architecting your highly available services. Prevention is always better than cure. Isn’t it! There is a pretty decent blog covering how to stop chkdsk from running on 2003 servers [including cluster servers] today we will see what options do we have for failover clustering 2008/R2. In failover clustering we have much more options of controlling the chkdsk behavior.

DiskRunChkDsk API--Determines whether the operating system runs chkdsk on a physical disk before attempting to mount the disk. Setting DiskRunChkDsk to FALSE causes the operating system to mount the disk without running chkdsk. With DiskRunChkDsk set to TRUE (the default), the operating system runs chkdsk first and, if errors are found, takes action based on the ConditionalMount property. The following table summarizes the interaction between DiskRunChkDsk and ConditionalMount.

The settings can be changed using the cluster.exe
cluster res “cluster disk ” /priv DiskRunChkDsk=value [see below for corresponding value, example 0 is default]

Chkdsk options

0 (Default): Run Normal Check. If corrupt, run chkdsk to fix the problem. Normal Check: Open the files in the root of the volume. Check volume dirty bit)

1 Run Verbose Check. If corrupt, run chkdsk to fix the problem. Verbose Check: Recursively open all files in the volume. Check volume dirty bit.

2 Run Normal Check. If corrupt, run chkdsk to fix the problem. If not corrupt, run chkdsk in read-only mode on the volume in parallel (i.e. online will proceed (and might complete) while chkdsk is running in read-only mode). We run chkdsk on a snapshot of the volume) and proceed with online.

3 Don’t do any File System check. Always run chkdsk on the volume.

4 Don’t do any File System check. Never run chkdsk, online disk without any File System check. Please note that this will also disable IsAlive/LooksAlive File System checks.

5 Run Verbose Check. If corrupt, fail online, don’t run chkdsk. User intervention required.

6 Suppresses volume creation/online/mounting during disk resource online. Disk is in offline read write mode, i.e. the disk is readable/writable using raw block level IOs. [ I doubt this is supported]

So let’s say as per standard maintenance task I need to run chkdsk every quarter or I need to take weekly backup in that case I can leverage the cluster maintenance mode to do the maintenance activity. Administrators use tools such as ChkDsk and VSS as part of weekly maintenance operation to ensure that disks are functional and there are no operational issues. These tools require exclusive access to the volume during their run. While these tools are in use, applications cannot read or write to the disk. The administrator expects the disk maintenance to succeed without ChkDsk failure and without a failover of the disk that the ChkDsk is run against. Under normal circumstances, the cluster disk resources will fail over when ChkDsk (fix error mode), VSS restore or any other tool that locks or dismounts the volume is run against a clustered disk. These tools fails part way through since cluster disk resource fails its health check that causes the cluster service to fail over the disk to the other node. This causes the node where these tools are run to lose access to the disk.

The following checks are performed on any disk that is online and that is managed by the cluster:

File system level checks: At the file system level, the Physical Disk resource type performs the following checks:
LooksAlive: By default, a brief check is performed every 5 seconds to verify that a disk is still available. The LooksAlive check determines whether a resource flag is set. This flag indicates that a device has failed. For example, a flag may indicate that periodic reservation has failed. The frequency of this check is user definable.
IsAlive: A complete check is performed every 60 seconds to verify that the disk and the file system, or systems, can be accessed. The IsAlive check effectively performs the same functionality as a dir command that you type at a command prompt. The frequency of this check is user definable.

Device level checks :At the device level, the Clusdisk.sys driver keeps checking every 3 seconds on PR table on LUN to make sure that only the owning node has ownership and can access that drive.

Maintenance mode is a mechanism provided through cluster.exe and the Failover Cluster API that places the specified resource in a mode that will disable health checking. After maintenance mode is enabled for a resource, Resource Monitor will ignore health check calls on the resource even though the resource is left in online mode. This will allow tools like ChkDsk to function against a resource that is in maintenance mode. Administrators should note that while ChkDsk is running the disk resource is not available to the application even though the resource is in online mode. When you put a disk in maintenance mode, this setting is an in-memory state and is not saved in the cluster registry hive. This change is not a persistent change. The next time that a disk is brought offline and then back online, the disk reverts to its standard behavior [we will see this later in article]
If there is any change to the state of the disk resource in maintenance mode, the maintenance mode setting is disabled. The maintenance mode setting is disabled when the following conditions are true:

 
Maintenance mode will remain on until one of the following occurs:
 
You turn it off. The node on which the resource is running restarts or loses communication with other nodes (which causes failover of all resources on that node). For a disk that is not in Cluster Shared Volumes, the disk resource goes offline or fails.

Ok fair enough…lots of talking..l picked one of my cluster shared volume “SR” and by default it is 0 [as seen below]. I am going to put this CSV disk resource in maintenance mode and then will run chkdsk on it. Ideally one should either save state or preferably properly shutdown all the VM ‘s who’s VHD are placed on this CSV before putting disk resource into maintenance mode. You will get a message that all dependent services and applications will be brought offline and cluster shared volume won’t be accessible from c:\clusterstaorage namespace.

C:\>cluster.exe res SR /priv
Listing private properties for 'SR':
T Resource Name Value
-- -------------------- ------------------------------ -----------------------
D SR DiskRunChkDsk 0 (0x0)


As I dint turned off my dependent VM’s before doing this action cluster service had to do it.
 


It also Removed access through the \ClusterStorage\volume path, however still allowing the owner node to access the volume through its identifier (GUID). This action also suspends direct IO from other nodes, allowing access only through the owner node. As I mentioned earlier “When you put a disk in maintenance mode, this setting is an in-memory state and is not saved in the cluster registry hive. This change is not a persistent change” we do not see this setting in registry and nor via cluster.exe output as 1 instead of 0 even after putting the disk in maintennance mode [though it’s very strange as then those properties should not be there or may be shoud be displayed in a different way] [see below]


 
 

Now once my CSV disk resource is in maintennance mode I need the GUID to run chkdsk on it. For a non CSV disk this GUID is not required as you will have a drive letter. I can either fetch the guid from mountvol.exe or powershell as shown below. though its more reliable and easier to take it from powershell if you have multiple CSV disk resources per node.

PS C:\Users\Administrator.UTOGWE> get-clustersharedvolume "SR" fc *
class ClusterSharedVolume
{
Name = SR
State = Online
OwnerNode =
class ClusterNode
{
Name = labw2k8hypv-1
State = Up
}
SharedVolumeInfo =
[
class ClusterSharedVolumeInfo
{
FaultState = 4
FriendlyVolumeName = C:\ClusterStorage\Volume1
Partition =
class ClusterDiskPartitionInfo
{
Name = \\?\Volume{ee372d39-9ec6-11de-b4ae-0017a4770008}
DriveLetter =
DriveLetterMask = 0
FileSystem = NTFS
FreeSpace = 618369548288
MaintenanceMode = True
RedirectedAccess = False
}
]
Id = 8e62ea00-9763-4c21-86d1-4a05708be24a
}

C:\Users\Administrator.UTOGWE>mountvol
Possible values for VolumeName along with current mount points are:
file:///?\Volume{94752fcf-46ad-11de-8463-806e6f6e6963}\
C:\
file:///?\Volume{ee372d39-9ec6-11de-b4ae-0017a4770008}\
*** NO MOUNT POINTS ***
file:///?\Volume{cd7b6371-b926-11de-85bb-0017a4770008}\
H:\
file:///?\Volume{cd7b6378-b926-11de-85bb-0017a4770008}\
I:\
file:///?\Volume{f4977a94-6176-4152-8a26-e3347856750d}\
F:\
file:///?\Volume{cd7b6384-b926-11de-85bb-0017a4770008}\
G:\

 
Now once I have the guid with me its piece of cake as I just need to run the command for defrag or chkdsk..and here we go.

C:\>chkdsk file:///?\Volume{ee372d39-9ec6-11de-b4ae-0017a4770008}\
The specified volume name does not have a mount point or drive letter.
C:\>chkdsk /f \\?\Volume{ee372d39-9ec6-11de-b4ae-0017a4770008}
The type of the file system is NTFS.
Volume label is SR.
CHKDSK is verifying files (stage 1 of 3)...
256 file records processed.
File verification completed.
0 large file records processed.
0 bad file records processed.
0 EA records processed.
0 reparse records processed.
CHKDSK is verifying indexes (stage 2 of 3)...
324 index entries processed.
Index verification completed.
0 unindexed files scanned.
0 unindexed files recovered.
CHKDSK is verifying security descriptors (stage 3 of 3)...
256 file SDs/SIDs processed.
Security descriptor verification completed.
34 data files processed.
Windows has checked the file system and found no problems.
786428927 KB total disk space.
182458068 KB in 43 files.
32 KB in 36 indexes.
0 KB in bad sectors.
90215 KB in use by the system.
65536 KB occupied by the log file.
603880612 KB available on disk.
4096 bytes in each allocation unit.
196607231 total allocation units on disk.
150970153 allocation units available on disk.

Well this is one way and then there is another easier way to do this via powershell. Another reason for you to fall in love with powershell.

PS C:\Users\Administrator.UTOGWE> get-help repair-clustersharedvolume
NAME
Repair-ClusterSharedVolume
SYNOPSIS
Run repair tools on a Cluster Shared Volume locally on a cluster node.
SYNTAX
Repair-ClusterSharedVolume -ChkDsk [-VolumeName] [-Parameters ] []

Repair-ClusterSharedVolume -Defrag [-VolumeName] [-Parameters ] []

DESCRIPTION
This cmdlet runs chkdsk.exe or defrag.exe on a CSV volume. It will turn maintenance on for the volume, move the cluster resource to the node running this cmdlet, run the tool, and then turn maintenance off for the volume. This cmdlet has to run locally on one of the cluster nodes. To run remotely, use PowerShell Remoting.


Once chkdsk/defrag has finished you need to turn off maintennance mode and manually turn on the virtual machines dependent on the CSV disk resource.


Hope this article would have given you an insight of what are the various disk maintennence options we have with failover clustering and how we can take right decisions during architecting our high availability solution.

Refferences
http://msdn.microsoft.com/en-us/library/aa370983(VS.85).aspx
Gaurav Anand

Green Storage - Need of 21st Century IT Infrastructure

The storage infrastructure within your data center will have just pushed up to 70%more carbon into the atmosphere, consumed up to 70% more power and up to 70% more cooling than it needed to. Over a year, a typical 42TB storage solution will push 8.9metric tons of CO2 into the atmosphere that otherwise could have been completely eliminated by a power efficient storage subsystem that meets or exceeds ALL the same performance, reliability and cost requirements demanded by your business.It’s easy to ‘tune out’ those kinds of talking points as IT professionals have grown more and more cynical of vendor marketing that always seems to over-promise and under deliver. But if green initiatives play a role in your organization’s priorities, power consumption solutions to the storage infrastructure are one of the easiest to implement and, thus, belong at the top of IT consideration.
Clearing the air on green storage

In an age of energy awareness, somehow the storage infrastructure within data centers has largely flown under the radar. Public awareness of ecological conservation is turning off lights, replacing incandescent bulbs, innovating greater levels of vehicle fuel efficiency, all while IT professionals continue to purchase and use the same growing amounts of storage that are consuming more power and pushing more carbon than they did ten years ago. In a world that has moved from incandescent to fluorescent, the vast majority of data centers are still using the same wasteful, storage systems that haven’t kept pace with power efficiency progress. Sure, vendors want to jump on the “green bandwagon” and claim “green storage” when, in fact, the only thing green about their storage is the color of the box it came in and the additional cost of the software you had to buy. It’s become harder and harder for IT professionals so see through the green smoke screen of vendor marketing.
Key Questions to Ask
1. Which vendors offer the most ecologically friendly storage solutions?
• Which vendors offer green solutions?
• What kind of solutions do they offer?
• How are they different from other vendor offerings?
2. How much of a reduction in power consumption and carbon production can be expected over a typical array?
• Are claims validated by lab reports?
• Can claims be substantiated by customers in real world scenarios?
3. Are there performance penalties to be expected in exchange for power efficiency?
• What kinds of applications are supported by green technology solutions?
• Can the vendor’s green storage technology be leveraged in SAS environment as well?
4. Do green storage technologies incur additional expense?
• Are there any additional costs incurred for green storage?
• Are there associated license fees?

The dirty secret is that some storage vendors feel justified to make a “green claim” when making the most minor of power efficiency improvements, e.g. a slightly improved power supply or the promise of a piece of software to utilize less storage which just costs you more money in the end. The way some companies try to stake a claim in the “go-green” trend is akin to a monster truck going green with recyclable seats.Storage vendors try to reduce the wattage of a fan while their disks needlessly spin at full speed when idle and call it a “green solution.” And of those who spin down, most deliver a green storage benefit that comes at the price of performance - a price not many applications can afford. The world of green storage marketing is so upside down that one storage system, which reduces power consumption by a meager 1%, can sit right next to another storage system that can reduce power consumption by a whopping 70%, and both are marketed as “green solutions.” More than ever, IT professionals have to look past “green claims” and inspect actual consumption reduction.
Green opportunity in Today’s Storage Infrastructure
The economic and environmental responsibility of our age is demanding more from disk storage vendors. A lot more. One of the reasons that storage energy waste has largely flown under the radar in the data center is because so much attention has been placed on the largest consumer of energy and capital expense in the data center — servers. However, with the advent of server virtualization and blade servers, IT professionals have made significant power improvements on the server level. Now that the server power problem is being addressed, attention has turned to the second largest consumer of power in the data center — storage. While servers may be the largest consumer of power in the data center, storage is not far behind accounting for 40% of all power consumption in the data center. With application servers becoming more efficient, it’s just a matter a time until storage becomes the top consumer of power, the largest producer of carbon and most significant source of waste in the data center. And with 50% aggregate data growth year over year, the power inefficiencies of today’s arrays can hardly be tolerated any longer from both an ecological and economic perspective.To understand the gravity of the problem, one must understand the power footprint of today’s data center. It is estimated that 1.5% of all the energy consumed in America comes from data centers, which is equivalent to the power consumption of 5.8 million households and exceeds to the total power output of all the coal power plants in the U.S. 40% of the power consumed by a data center comes from its storage infrastructure. A single watt saved on the drive level does more than just save power consumed by the drive; it ripples throughout the entire cooling infrastructure, power distribution infrastructure and ultimately slashes the carbon production from all three sources. For every watt saved on the drive level, roughly 3 watts end up being saved at the meter.2

Four Steps to Greener Storage Infrastructure:
1. Utilize power efficient storage arrays
2. Increase existing storage utilization with virtualization and thin provisioning?
3. Reduce storage with deduplication and compression
4. Consolidate data to more power efficient tiers


Carbon Footprint of Today’s Storage Arrays
Ecologically, since the industrial revolution, increased amounts of greenhouse gases have been emitted into the atmosphere — dramatic increases in CO2, methane, tropospheric ozone, CFCs, and nitrous oxide. The concentration of CO2 alone has increased by 36% since the mid-1700s. These levels are considerably higher than at any time during the last 650,000 years — the period for which reliable data has been
extracted from ice cores. Less direct geological evidence indicates that CO2 values this high were last seen approximately 20 million years ago. Fossil fuel burning has produced approximately three-quarters of the increase in CO2 from human activity over the past 20 years. The remainder is due to land-use change — deforestation in particular. The issue of climate change has sparked debate about the benefits of limiting industrial emissions of greenhouse gases verses the costs that such changes would entail. EPA Administrator, Lisa Jackson, announced in Copenhagen that the agency had finalized its finding that greenhouse gases, including carbon dioxide, pose a threat to human health and welfare. The EPA(US Envt. Protection Agency) will soon begin regulating greenhouse-gas emissions from power plants, factories and major industrial polluters. Data center regulation is only a matter of time. In the U.S., The House has already passed a bill that would cap U.S. carbon emissions at 17% below 2005 levels by 2020. The Senate is considering similar legislation. While global warming is not solved by any single action, the balance is dependent upon the cumulative effect of everyone doing their part. As individuals, the responsibility trickles down to things as simple as turning off a light or moving to a high-efficiency bulb to decrease one’s carbon footprint. In the data center, the problem is drastically larger, but, in many ways, very simple to solve.
Green Storage Best Practices
A variety of best practices can help us better understand efficiency. In storage, there are three things to consider to improve energy efficiency:
*The additional energy consumed because of inefficient devices
*The additional capacity required because of inefficient management
*The additional floor space required because of inefficient packaging
Conclusion
With the convergence of our current ecological challenge, we are all faced with our own individual responsibility. No single action can solve all of the problems we face today, but we can’t ignore that the best solution lies in the accumulation of many small changes.

Source:
1Report to Congress on Server and Data Center Energy Efficiency
Public Law 109-431. U.S. Environmental Protection Agency ENERGY STAR Program
2Energy Logic: Calculating and Prioritizing Your Data Center IT Efficiency Actions,
Emerson Network Power
www.nexsan.com

Tuesday, March 30, 2010

How configuring storage with Volume GUID works in Failover clustering

There is couple of new features and architectural changes in failover clustering and I am going to talk about new functionality of using Volume GUIDs instead of drive letters in this talk. You are recommended to apply this patch for this increased functionality.

951308 Increased functionality and virtual machine control in the Windows Server 2008 Failover Cluster Management console for the Hyper-V role [This is not required for server 2008 R2]

http://support.microsoft.com/default.aspx?scid=kb;EN-US;951308

Let’s see how this works behind the scene, so let’s dig in to find out what happens in the background. I have a 2 node Hyper-v cluster with Node and file share majority as the quorum model. To investigate this Guid behavior I presented new storage disk to my cluster and noticed the view in disk management, Failover cluster admin console, {HKEY_LOCAL_MACHINE\SYSTEM\MountedDevices} mounted devices registry key and mountvol.exe output. Node 1 is the current owner of this disk. On Node 1, we can observe that the newly presented disk shows up as disk 4 in disk management.msc with drive letter H:\ but in failover cluster admin console it appears as cluster disk 5, so there is no co-relation between two and we should not be confused about it. (See fig 1)


On node 2 the disk got the next available drive letter which is F:\ and got its own unique Volume GUID. So we see both nodes have got different local Volume GUIDs for same disk .Till now it is all expected as every machine should have its own GUID for a disk and they picked the next available drive letter on node 2 so all looks good so far (see fig 2)

FIG 1


                                                                 FIG 2

Now I went ahead and removed the drive letter so that we can use this disk with GUID (see fig 3)

FIG 3

This is what I observed (see fig 4) in failover cluster admin console…its displaying Cluster Volume GUID for the disk which is same as the disk’s Local Volume GUID of node 1 (remember node 1 owns the disk). On node 1 in mountvol.exe output we see a fresh new entry with modifier “no mount points” which was not there earlier. This shows that this disk is not using a drive letter now. On node 2 we don’t see any change on either mountvol.exe output or mounted devices registry key as seen in fig 5.
FIG 5

Now I opened Hyper-V console on Node 1 and created a virtual machine named “Guid behavior” as shown in fig 6. Here we provide the path using Volume GUID. As I am going to create a highly available Virtual machine later on, I will use the Volume GUID displayed in Failover cluster admin console for cluster disk 5 which is the recommended way of creating a highly available VM. (We will come to know why this is recommended way later in the blog)


FIG 6

However the question arises that the Volume GUID displayed in failover cluster admin console is same as local Volume GUID on node 1 however we have no reference of this Volume GUID on node 2. So when this highly available VM will move or failover to node 2 what will happen? How will it come online as node 2 has no idea of this Volume GUID showing up in failover cluster admin console…let’s see what happen.

I created a highly available virtual machine service and this is what I observe on node 1 and failover cluster admin console (see fig 7) which is all expected.
 
FIG 7

Now lets see what do we see on node 2 . we still see the exact same information (see fig 8) in mounted devices registry key and mountvol.exe output as seen before. No changes till now. So lets try to move this VM from node 1 to node 2 and lets see what happens….we see that in mounted devices registry key we now see 2 GUIDs for disk with signature 6fd2cfff, one is Local volume GUID which was orginally present on the node 2 before moving VM and other one is Cluster Volume GUID (Cluster replicated actually) which is being displayed in failover cluster admin console and is also the Local Volume GUID of disk with signature 6fd2cfff on node 1. (see fig 9).

So whoever node owns the disk before it was used as highly available disk in cluster will replicate the Volume GUID to all the remaining nodes and then that volume GUID is used by cluster and is known as cluster volume GUID. We will still see the Local Volume GUID on all the nodes along with cluster Volume GUID. You can distinguish between Local Volume GUID and Cluster Volume GUID easily by seeing the Mountvol.exe output as the Cluster Volume GUID will not show up on Node 2 mountvol.exe output however it does show up on node 1.

FIG 8


FIG 9

Well, this completes our discussion for the day and we now know how we can use GUIDs instead of drive letters for configuring storage and how cluster replicates these Volume GUIDs. There is one more word of caution that I have noticed and I will like to share. We can place virtual machine either via Hyper-V or system center virtual machine manager. Now using system center virtual machine manager I placed newly created virtual machine “Guid behavior” on node 1. While creating this virtual machine we have to provide storage for placing .vhd file and as you can see in fig 10, SCVMM picks the Volume GUID information automatically. There are scenarios when it may pick the local volume GUID instead of Cluster Volume GUID and that’s why we recommend that if you are creating highly available VM, make sure that you confirm the GUID used, is the GUID displayed in failover cluster administrator console.


                                                                FIG 10

This ends our discussion for today and I hope that you enjoyed reading the blog. Please come back on our  blog where we try to share information and Thanks for your valuable time. There is already a very nice blog written by Chuck Timon on this matter and I strongly advise you to read that too.

Configuring Storage Using Volume GUIDs in Hyper-V

http://blogs.technet.com/askcore/archive/2008/10/29/configuring-storage-using-volume-guids-in-hyper-v.aspx

GAURAV ANAND

Saturday, March 20, 2010

Storage Architecture changes for failover clustering 2008/R2-How Persistent Reservation works.

Last month I decided to do a failover cluster blog series and here I am doing the first one. Via this blog I will try to throw some light on the storage architecture changes and requirements of failover clustering. As we are aware, only storage that supports scsi 3 persistent reservations will be supported in failover clustering. Parallel scsi based storage is being deprecated and won’t be supported. There is a nice blog to read more why this might have been done.  The good thing is that due to this change now we no longer use scsi bus resets which can be disruptive on a SAN. In this article we will see what happens in the storage stack and how persistent reservation works for a physical disk resource and later a cluster shared volume physical disk resource.

                                                                                 Fig 1

We will take the case of physical disk resource first and then cluster shared volume later. Clusdisk.sys which has been modified with 2 functions seen in the figure 1 issue persistent reservation [PR] to the class driver disk.sys which flows down to MPIO.sys and vendor based DSM.sys or inbuilt MSdsm.sys [msdsm.sys is inbuilt device specific module provided with 2008 R2 server] The vendor based DSM driver is responsible for registering the PR on the storage object. The storage object maintains a registration table which contains entry from all the multiple paths available from multiple nodes. The registration table contains registration and reservation entry for all paths available from all nodes. We take an example of a 2 node cluster with dual HBA and hence we see that every interface from both nodes will have to register in the table with a unique key for each node. The rule is that you cannot register and reserve at a same time though there are some exceptions to that rule and we will discuss those later. This key is a 8 byte key unique to every cluster node. The low 32 bits of this key contain a mask that can be used to identify a Persistent Reservation that was placed specifically by a Microsoft Failover Cluster.  So assuming node 1 HBA1 is right now owning the physical disk resource you will see both registration and reservation entry  for HBA1 interface. In case HBA1 fails then depending on the vendor DSM configuration storage stack will automatically start using HBA2 without any end user intervention or failover of the disk resource. Every 3 seconds node 1 keeps coming back and checking the registration table. Nodes defend their reservations (every 3 secs…but is configurable and this setting might help you in troubleshooting some issues also) . During a split brain scenario challenger node 2 will come and enter a registration entry in the table but as per rule node 2 cannot register at same time. It needs to wait for at least 6 seconds before it comes back and enters its registration entries and own the ownership for the storage object. Meanwhile defender node 1 comes back after every 3 second and  sees registration entries from 2nd node  and it will scrub those entries. Challenging node 2 will come back after 6 seconds but will not be able to add reservation entries in table as its registration entries have already been scrubbed by node 1. This is the process of a successful defense. Now there may be a legitimate case where node 1 both interfaces are unable to access the storage or node 1 has rebooted or blue screened and in that case node 2 should win the arbitration of storage object. Let’s see what will happen in that case. Node 1 HBA1 is right now owning the physical disk resource and you will see both registration and reservation entry  for HBA1 interface.  Node 1 blue screened, so this time when node 2 comes to register and reserve it will successfully register and 6 seconds later on revisit it will put a reservation entry into table and will own the storage object. This is how the storage arbitration and scsi 3 persistent reservation works. You can see in detail the same process happening during the cluster validation test –validate scsi 3 persistent reservation as seen below in figure 2.

                                                                             Fig 2

Validate SCSI-3 Persistent Reservation
Validate that storage supports the SCSI-3 Persistent Reservation commands.
Validating Cluster Disk 0 for Persistent Reservation support
Registering PR key for cluster disk 0 from node Node 1
Putting PR reserve on cluster disk 0 from node Node 1
Attempting to read PR on cluster disk 0 from node Node 1.
Attempting to preempt PR on cluster disk 0 from unregistered node Node 2. Expecting to fail
Registering PR key for cluster disk 0 from node Node 2
Putting PR reserve on cluster disk 0 from node Node 2
Unregistering PR key for cluster disk 0 from node Node 2
Trying to write to sector 11 on cluster disk 0 from node Node 1
Trying to read sector 11 on cluster disk 0 from node Node 1
Attempting to read drive layout of Cluster disk 0 from node Node 1 while the disk has PR on it
Trying to read sector 11 on cluster disk 0 from node Node 2
Attempting to read drive layout of Cluster disk 0 from node Node 2 while the disk has PR on it
Trying to write to sector 11 on cluster disk 0 from node Node 2
Registering PR key for cluster disk 0 from node Node 2
Trying to write to sector 11 on cluster disk 0 from node Node 2
Trying to read sector 11 on cluster disk 0 from node Node 2
Unregistering PR key for cluster disk 0 from node Node 2
Releasing PR reserve on cluster disk 0 from node Node 1
Attempting to read PR on cluster disk 0 from node Node 1.
Unregistering PR key for cluster disk 0 from node Node 1
Registering PR key for cluster disk 0 from node Node 2
Putting PR reserve on cluster disk 0 from node Node 2
Attempting to read PR on cluster disk 0 from node Node 2.
Attempting to preempt PR on cluster disk 0 from unregistered node Node 1. Expecting to fail
Registering PR key for cluster disk 0 from node Node 1
Putting PR reserve on cluster disk 0 from node Node 1
Unregistering PR key for cluster disk 0 from node Node 1
Trying to write to sector 11 on cluster disk 0 from node Node 2
Trying to read sector 11 on cluster disk 0 from node Node 1
Attempting to read drive layout of Cluster disk 0 from node Node 1 while the disk has PR on it
Trying to read sector 11 on cluster disk 0 from node Node 2
Attempting to read drive layout of Cluster disk 0 from node Node 2 while the disk has PR on it
Trying to write to sector 11 on cluster disk 0 from node Node 1
Registering PR key for cluster disk 0 from node Node 1
Trying to write to sector 11 on cluster disk 0 from node Node 1
Trying to read sector 11 on cluster disk 0 from node Node 1
Unregistering PR key for cluster disk 0 from node Node 1
Releasing PR reserve on cluster disk 0 from node Node 2
Attempting to read PR on cluster disk 0 from node Node 2.
Unregistering PR key for cluster disk 0 from node Node 2
Cluster Disk 0 supports Persistent Reservation

But there is an exception –remember –[ The rule is that you cannot register and reserve at a same time though there are some exceptions to that rule and we will discuss those later.] Ok let us see where we use this exception. There were various changes that failover cluster team brought in storage architecture and one is that we do not arbitrate the same way as we used to do for MSCS clustering. We now never arbitrate for physical disk resources in case of a controlled manual movement of resources across node [manual move group process or quick migration] Though we always arbitrate for witness disk resources.  We only arbitrate for other disk resources in failure conditions. We use a fast path algorithm for such a scenario and why we do this! Because this is was a requirement for effectively supporting Hyper-v Virtual machines as resources in failover clustering. When we move a virtual machine resource from node 1 to node 2 the storage object containing the .VHD file needs to move quickly and cannot wait for 6 seconds for arbitration to take place [in case of a manual move group process or quick migration]. Leveraging Fast path algorithm fixes this challenge for us. Using fast path algorithm physical disk resource, in use by virtual machine group can move across nodes in less than 1 second approx or even lesser time.

So lets see what happens when I manually move a physical disk resource which is currently owned by node 1 i.e. node 1 has entry in the registration table for the storage object. Node 2 will wipe the registration table and scrub all the entries from all the paths from all nodes. It will place new registration and after successful registration within milli seconds will put a reservation entry in the storage object PR table which means that it owns the physical disk resource now and can bring it online. In such a scenario node 2 will not wait for node 1 for 6 seconds and this complete operations takes in approx less than 1 second. This is what happens in case of quick migration and when you manually move a group containing physical disk resource. However as said earlier in all circumstances we arbitrate for witness disk resources.





We all know that we can put a physical disk resource in maintenance mode to run chkdsk [for exclusive access]or other maintenance operations. So what happens when we put a physical disk resource in maintenance mode! All the non owner cluster nodes [who have their entry in registration table] will not be able to access the storage object as disk resource will be fenced from all of them and then owner cluster node will also remove its persistent reservation from the storage object PR table. So in other words we temporarily made this storage object a non clustered storage object enabling chkdsk or other Maintenance operations to run on it.

Now lets jump on cluster shared volumes in failover clustering 2008 R2. Cluster shared volumes is a feature of windows server 2008 R2 clustering which allows the different nodes of cluster to have concurrent access to the LUN where highly available virtual machine's VHD is stored. CSV allows multiple VHD per LUN and removes the traditional one VM per LUN issue. Earlier clustered virtual machines can only fail over independently if each virtual machine has its own LUN, which makes the management of LUNs and clustered virtual machines more difficult. As of now CSV feature is supported for use with windows server 2008 R2 hyper-v role only. As its obvious that there is no arbitration happens in case of CSV physical disk resource so how does all the nodes access the storage object at same time. Again we have same concept of PR table on storage object and we take an example of 2 node cluster. Node 1 is the coordinator node in this example which means that node 1 owns the CSV resource. Node 1 will have  a unique key in the reservation table which grants it access  as owner node and all the ntfs metadata writes are achieved via this node. So if there is a need to modify ntfs metadata of this CSV physical disk resource from node 2 it will be passed to node 1 and then node 1 will take care of it. However node 2 can still read/write to the CSV physical disk resource. The difference here is that non coordinator nodes will have their read/write key into the registration table always while coordinator node will have a unique key in the reservation table of the storage object.

So how can I remove a PR manually? Cluster.exe [node-name] /CLEAR[PR]:device-number is the command to clear a persistent reservation manually. You might have to use this command during troubleshooting of PR issues. You can also change the disk arbitration interval from 3 second to customized value but remember that there will be repercussions of modifying this resource private property and needs proper testing from storage vendor and remember that this setting comes in picture only in case of failure conditions and not for manually controlled operations.

I hope this article would have given you an insight into how persistent reservations works for failover clustering for physical disk resource and cluster shared volume disk resource and with this understanding you can do more effective troubleshooting & planning of storage related clustering issues. Thanks for your time and hope this is helpful.

GAURAV ANAND

The information based is "AS IS" and based on my personal understanding of the Technology.