Saturday, August 7, 2010

reason number 9039808204802 why i hate RPM

[root@dhcp9001 ~]# rpm -Uvh dhcp-3.1.9999-2.cbs.i386.rpm
########################################### [100%]
package dhcp-3.0.5-7.el5 (which is newer than dhcp-3.1.9999-2.cbs) is already installed

NO, IT IS NOT NEWER. FUCK YOU. INSTALLLLLLL!!!!!

Tuesday, July 27, 2010

quality control in your network

(disclaimer: i've never worked in quality control, but this is my view of it as someone who has had to work with QC)

while i sit here at 3 in the morning waiting for a server daemon to dutifully seg fault leaving me to continue debugging, i reflect on how quality control is lacking from so many networks both large and small. a single oversight *can* mean the difference between your company losing money or going under, so you must be aware of any potential problems at all times.

first of all, what is quality control and why do you need it? well, chances are if you do your job halfway right, you're already doing it. quality control is basically double-checking that the methods you use to do your job are correct. it doesn't verify that the final product is good; it's more like, tripwire in procedure form. it's making sure that things work the way you expect them to. you are already doing it when you verify what development libraries are installed. when you run unit tests against your software. when your change management system verifies a user is allowed to commit that particular piece of code, or restart that service. it's checking the tapes to make sure the backup robot is functioning. it's verifying configs are written properly on the router and the updated ones are regularly saved in version control.

typically you don't need much quality control in the average network. some product development may require strict control and observation of policies and procedures, which is usually only reinforced due to the risk of random audits or inspections. depending on your environment you may be required to do very little or no quality control at all. but i'd like to tell you about the quality control you should be doing.

the quality control people aren't usually technical people. a lot of the time they'll work with a team member of whatever they're checking out, ask questions and make notes. the first big formal procedures don't include everything. usually details get hashed out while the QC engineer talks to someone (a dev in this case) about what they do and how to check that what they did worked.

the basic principles you should keep in mind when applying quality control to your network are as follows:
1. Keep It Simple, Stupid. it doesn't have to be verbose or complex. be flexible. be easy.
2. it should be possible for someone to check the work of the quality control engineer(s).
3. you don't want to define how everything works; only how to tell if it's working as expected.
4. your goal is to make sure there are no catastrophic failures. you don't have to account for every blip along the road as long as the road is open.
5. start with the big things and move down once all the big things are covered. close off those single points of failure and move on to the other pressing issues.

hopefully this post will help to give you an idea of how you can apply quality control to your network to get an improvement in the overall quality of service you provide. half of this is just making sure things works right, and then the other half is reviewing that there are records that it's been used properly in the past. here's some stuff you can do.

developers
double-check that your software is being created correctly. check that the libraries on your development boxes match up with what's going into QA or production. use unit tests on your code. make sure everything goes into version control *before* it ever hits QA or production, and make sure you know who made what change and why. make sure the method of deployment can be reversed at any time. make sure you follow change management procedures when necessary.

sysadmins
double-check that you've confirmed with everyone before you push a new piece of software to QA or production, and that you can roll it back when necessary. so check that your change management is working. it's good to have a list of the major and sub-major software that different development teams rely on (usually libraries) and get a change-management approval before ever pushing this stuff out. do it early so devs have time to test their shit with the new software. make sure your backups work correctly. you should be able to confirm logs and destination files to ensure the backups are going well regularly. if those or any other automated process fails it should generate an alert, and you should be able to verify those alerts are going out as expected (did /var fill up and is sendmail unable to work now?). make sure all security patches are applied in a timely manner. make sure all service-monitoring systems are working, and that failover of critial systems is in place and works as expected. make a list of all critical infrastructure and make sure all of it has hot-spare failover systems waiting in the wings. provide for methods to remote troubleshoot in the event of total internet or system collapse. make sure any network gear you depend on also has hot failovers that work.

there's more implementation-specific details you sometimes need to get into with QC. i want to get more into how to begin making procedures for these systems but it's way past my bedtime. will continue when i am not so sleepy.

Tuesday, June 29, 2010

you poor apple users

i really feel sorry for people who recently got the iphone 4. i sat at work, listening to my co-worker half-heartedly explain how his iphone sometimes gets really bad reception, but that it's ok. that he naturally holds his phone in exactly the way necessary to avoid the signal loss. how it's not alright that it has this problem, but it's also ok because it doesn't affect him for the most part.

i just had this twang of empathy. like, i see now... you really don't have a choice. what are you going to do, return the phone and get a non-iphone 4? it's not an option for him. in one sense because it's now ingrained into his life; the apps, the services, the way the phone feels, the way he uses it.... that's part of him now. he can't let it go. it's scary to think of not having an iphone. and that's the second sad thing. it's actually changed his manner of thinking and now he can't get away from the thing. it's like an addiction. he can't see himself without it, and now he's trapped by it.

this is a guy who less than a year ago had never owned an iphone. he's technically a "late adopter." yet in less than a year apple has not only converted him, they've made him their slave. in some ways i can relate... since getting a smart phone i feel like i need to have a more powerful one. i need to be able to stream video, or use irc or ssh remotely, or browse web pages fast and in full render. do i actually need these things? fuck no. i was perfectly fine with my old brick sony ericsson with the wonderful camera, sending pictures and using web pages just fine. now i sit in at&t stores trying to sell myself on buying the newest piece of shit android phone which still doesn't have as good a camera (and definitely no xenon flash) like my old phone.

so i get it now. i'll stop pointing fingers. i'll stop trying to convince you. because i know you can't escape it. you're trapped in a tar pit of technology and you can't get out even if you wanted to. i feel sorry that you're stuck with a (somewhat) shitty phone with a shitty app development model on a really shitty carrier. i wish we could all just have open-brand open-carrier open-market smartphones and share information and use the internet as freely on our mobile devices as we can on our computers.

but we can't. we probably won't ever, or not for a long time anyway. they figured out the way to trap us and steal more of our money than they ever did with the PC or laptop. it's overly-expensive unlocked phones and contracts and insurances and data plans and messaging plans. i pay $110 a month for a "standard" phone plan with "unlimited" data and "unlimited" messaging. that's $2,640 dollars every 2 years (without the extra amount i pay for a new phone every 1.5 years). i don't even buy a laptop or PC that often, or gaming machines or games. i just paid $520 for an almost-new laptop, the first big purchase in over 5 years.

this is kind of depressing me. i feel like i'm trapped too.

Monday, June 21, 2010

what's wrong with my breakfast?

a recent survey found that 85% of a selection of engineers don't use twitter. they cited not caring what people had for breakfast as a reason they don't use it. this is my response.

what's so wrong with my breakfast that you can't stomach the information? i realize it's useless. i realize you don't care. but also realize, i don't care that you don't care. i am one of the mindless drones of twitter and [formerly] brightkite and facebook and myspace that update our status with whatever mindless drivel we happen to think is important in the moment at that time.

there's not much logic involved. wanna say something? say it. people will listen or they won't. but i don't expect them to. it's more of a general smattering of my thoughts and a few choice life experiences that people can refer to if they wish in the future. it can be a way for a potential employer to see if i might be a fit for their organization. it might be a way for a single young lady (or lad?) to determine if i'm worth sending a poke/message/direct tweet/whatever. perhaps i just want people to know i know certain things, or have certain opinions. whatever it is it's certainly one part exhibitionism, one part honest sharing of experiences and thoughts.

it would be nice if it had more function. perhaps a symbol prefix to be flagged in a number of ways: exclamation mark for an urgent or important message, question mark to "crowdsource" (ugh), pound sign to advertise an event, dollar sign to advertise a neat deal or other ad. extra modifiers could be postfixed to give further detail about the post. then each user could set their preferences of what kind of information comes into their stream from their friends.

that doesn't exist, though. what does exist is "i'm at place XYZ and i'm having a blast!" or, "this is a picture of a fucking AWESOME fish taco" or, "who else thought iron man 2 was kind of lame?" my posts aren't intended for the general public. they're not always even meant to be useful. i send this shit out because i'm bored and i want to share with my friends. they don't even have to interact with me. in some small way i am enriching lives with my pointless drivel. i help kill boredom. sometimes i share articles which i find useful or informative. sometimes i share music that means something to me.

you probably won't find it useful at first. you'll probably be too bored with it to even start. but once you have a good chunk of friends added into your account, you'll notice you can interact with them. you can follow what some of them are doing. maybe even find out something you wouldn't have without the service. yes, we should all have real lives and not be so connected to a text interface all the time. but sometimes technology can be more than just a distraction.

granted: most of twitter is pointless, and the examples i give which contain somewhat useful information is by and large not close to the majority of the tweets out there. twitter kinda sucks. but facebook is a good example of what it could be. the old brightkite was a good example of what it could be. maybe in the future all these techie news aggregating sites can turn into an honest-to-goodness social network of just nerds so we can all collaborate interactively and enrich each other's work and personal lives on common ground, where we can hack on it and do what we want with our own little home on the internets. but that's probably just a pipe dream.

i don't care if you use it because i don't watch my twitter feed tbh. but you can get a PixelPipe account, add your friends that are on the various social networks, and share with them all at once. you don't even have to follow what they're saying. just say something. try to make it useful; who knows, maybe you'll start a trend.

Wednesday, May 19, 2010

linode: a paradigm of indifference

my linode was unavailable for 2 hours yesterday from 11:50am to 1:47pm. the linode itself stayed up but its connection to the internet was down. the newark datacenter it's hosted in went down and though i could apparently console it, i could not connect to it outside the control panel linode has for its users to admin their node. obviously they had some kind of connection established between their different datacenters if i could console it.

so i inquired if i could migrate the node to another datacenter.... not while the internet was down, apparently. and i couldn't do it without submitting a ticket and waiting for a response. to me this is aggravating. why is it open-source and commercial VPS solutions can all automatically migrate nodes from one dom0 to another but linode cannot do it automatically? a question that will probably never be answered.

what troubled me more than the downtime and lack of actual information or an ETA was the response from IRC support channels. lots of people were joining the channel, asking if there was a problem, what the cause was, and when it would be resolved. for a service we all pay a minimum of about $24 a month, probably the highest amount you can pay for a VPS with the same specs, this should be a completely acceptable set of queries. some people pay a lot more than that, btw. you'd think they would want to be courteous and attentive to all their questions, right? or, i dunno, put in the fucking topic of the channel the current status?

they didn't update the topic. most of the people coming in were treated with sarcasm, and no effort was made to silence those who were in the channel merely to be a dick - which is common on irc. but not a good idea for an official support channel for a paid service. most people who were there with concern that their services were unavailable were treated with a general indifference and in many cases as if they were simply pests and whiners. i guess most customers figured a 2 hour downtime in the middle of the day was something worth complaining about. i guess we were wrong.

there is still no explanation for the issue on the linode status page. at first we were told in the channel it was an issue with level 3, and the london datacenter was also down. then it was only the newark datacenter. through continuously attempting console access and traceroute/tracepath i saw how the connection would drop right at the router before my linode is usually hit. eventually the whole newark datacenter's routes stopped, but soon came back; sometime during this time console access was terminated, so clearly there was some real traffic going to linode inside the DC throughout.

because the datacenter's routers were inaccessible via traceroute for a short time i assume they were somehow convinced there may indeed be a problem with their routers, and so i'm not ready to completely blame linode. but certainly they had some connectivity and could have offered to migrate our nodes to another DC so, for example, we may have only had a 1-hour downtime instead of 2 hours. no such attempt was made. the estimated refund for this downtime is something like $0.08, which is of course nothing compared to the amount lost in business to the linode users who suffered the downtime (SLA's never reimburse near the amount you lose, so you shouldn't count on them that much). not that it would have made a difference, but the heartless way we as customers were dealt with makes me really dislike this company. now i know if i have a problem in the future, nobody's going to really try to help me. apparently they don't need me as a customer. and that's OK; i can find cheaper hosts with bigger caps elsewhere.

in the end this will be helpful. i already had a secondary VPS i paid about $5 for monthly, whose billing i let lapse out of laziness. this event will help motivate me to move back to that host and a couple others and have truly redundant services for the same cost as the one node i'd been paying for at linode. sure their web interface is fancy and you have a good deal of freedom. but considering the availability and bad customer service? i think i'll go with the cheap guys.

if you want a cheap VPS, check out Special VPS and Low End Box. they review and give promo codes for low-end VPS providers. by reading their reviews you can learn how to spot shady and unreliable hosters. do your research!

Thursday, April 8, 2010

how to make a product everyone will buy

  1. Make a hardware platform only you control and manufacture. Also make sure it looks very pretty and is reliable.
  2. Make an operating system that's user-friendly, simple, and very pretty with some killer apps and an easy dev environment. But make sure it only works on your hardware.
  3. Make software to provide all the personal needs one has with a computer and make it all tie together seamlessly. Make sure it's user-friendly, simple, and very pretty.
  4. Make accessory products which provide for things people want day to day, make it work with your hardware and software, and tie it all together seamlessly. Also, make sure it's very pretty and reliable.


Now release anything and make sure it ties together with all the previous products seamlessly, is user-friendly, simple, reliable, and very pretty. It doesn't even matter if it has a purpose or is redundant: people will buy it. It helps if you have the world's greatest PR/hype machine and if you can make people believe they're superior to someone else by owning these products. Above all make sure it is always very pretty. In this vein the product is like a luxury car: completely impractical and unnecessary, but people pay a premium for something that looks fancy and probably doesn't provide any benefit over a cheaper less pretty device.

Monday, March 15, 2010

Better Security [tips]

You've got network intrusion detection and stateful firewalls. Your kernels are patched as far as they can go for exploit prevention. You're using OpenSSH. That's awesome. Now why is it someone can still penetrate your precious servers so easily?

When you begin to secure something (anything, really - buildings, documents, servers) you have to consider everything. Each factor which could possibly be targeted in an attack could be used with any other factor to increase the likelihood of a successful compromise. So each factor has to be looked at in conjunction with every other factor. Yes, this is usually incredibly tedious and mind-bogglingly complex. To help mitigate this you can design preventative measures around each possible attack vector. In other words, add security to everything.

In the example above there's loads of attack vectors just waiting to be leveraged. One example is OpenSSH. A lot of people just use it in it's default form and never add any security to it. This will lead to an exploit. If you allow password entry to an OpenSSH server, just assume it's been compromised. It's so easy to observe a password being typed or intercept it somewhere else it's laughable. Not to mention people hiding passwords under their keyboards or on their monitors! No, a password-protected SSH key is the minimum you should use to allow access to a server. The "something you know, something you have" two factor authentication is far more secure than a single factor. I should stress that this is only true when properly implemented, as bad two-factor can be even less secure than strong one-factor. For more on authentication factors read this and take note of the ways different factors can be exploited (don't rely on just biometrics!).

In newer versions of OpenSSH there are even more methods to harden the authentication process, such as certificate authorities and key revocation lists. Also disabling root logins, having a set list of users allowed to authenticate, disabling old deprecated protocols, ciphers and algorithms, and explicitly dropping any connection with conflicting host keys is a good idea. You should even consider the libraries used by the application - were they built with buffer overflow protection? Is PAM enabled? One need only look around to see the underlying systems of your very critical remote administration software could be rife with potential exploits. For every one exploit known there is probably ten unknown waiting to be found.

Now think about what you have access to on your own system. Consider for a moment what would happen if an attacker used the same methods you do to gain access. Would it be difficult or easy for them? If it's easy for you to access the system, it may be for them. Try to make it more difficult even for yourself to gain access and a potential attacker will have a hell of a time trying to leverage something you've left unguarded. Make your firewalls incredibly verbose and restrictive; you'd be amazed how little can be done to a system when an attacker doesn't know exactly how to use it. Require multiple levels of logins before root can be obtained, and try to minimize any need to get to the root account.

Make all of your services run as unprivileged users. Make scripts to be executed by sudo which don't take any options (and clean their environments) that take care of tasks you may need root to perform. Any admin should be able to perform basic administrative tasks as a non-root user. Make all services controllable by an "admin" group, with each service having its own unique user to minimize attacks from one service to the next. Most services can be configured to start up and bind to a privileged port and drop to an unprivileged user, but for those that cannot there are methods (SELinux, etc) to work around restrictions in an application or system.

Make your configuration also be applied by a non-root user. A good way to think of configuration management is "let the service manage itself." Create your base model or template of configuration management scripts and then create service-specific configuration that can be run by the user of that service. In this way you don't need to worry about an attacker pilfering your configuration management and applying rules to all the machines as root. You can also create more fine-grained controls in terms of what admin (or group) can configure what service. You don't need to worry about a "trusted" user compromising your whole network if you only explicitly grant them access to the things they need to manage.

In fact, consider time-based access control for your entire network. You should expire SSH keys and user access for different services around the same time you expire old passwords. This will force you to improve the method users have to request access and hopefully increase productivity and responsiveness in this area of support. Just don't fall into the trap of allowing anything anyone asks for. Make it easy to get their manager to sign off on a request so at least there's some accountability; you can only benefit in terms of security if somebody thinks they might get fired for granting access willy-nilly.