2012-02-05

Abusing the system?

I blogged about abuse of MSO texts.
I blogged about people abusing ADR.

Did I ever expect to be woken by someone with an MSO text over an accounting error - paying me £3K by accident, and urgently needing it sent back so he can pay the right person. It wakes quite a few staff as it happens. It is about as far from a Major Service Outage as you can get.

And then apparently threatening that he could go to ADR* and cost me £350, twice, if I did not help. Vexatious? I very nearly ceased his bloody line on the spot when he said that.

I have sent the money back, before midnight (which was apparently his deadline to pay the right person).

I am at a loss for words to be honest. I don't know what to say. I am going back to bed now.

P.S. they are now talking on irc about using MSO texts to request pizza and taxis...

*The exact line was:-
23:35 < ydorg> having two ADR's could cost 700 quid

2012-02-04

Snow

Someone on irc says it is snowing. I say it is not. But to check I check the weather app on my iPad and look up where I am and it says "Light snow". Oh, apparently it is snowing.

Actually getting up and looking out of the window if very much an after thought, and it confirms that there is indeed light snow here.

What has the world come to? If the weather app and the window had disagreed - which would I have believed? :-)

2012-02-03

Sharps box

It is a bit of a sad day when you have to get one of these, but given that my mother has been diabetic for nearly 50 years now it is not surprising I now have one.

I also have a large box of needles and some "pens" of insulin. They are quite cunning - disposable pens with 300 units of insulin pre-filled. Dial up the dose, fit a disposable needle, close your eyes, grit your teeth, and push the button... Something like that anyway...

I get my training on this next week when I start daily injections.

Ho hum...

ADR

I have to post something - I am so stressed over this - if I post on here I will feel I have parked the problem, at least for now. Sorry it is long and boring...

This should be a case for putting on our web site as a major success story. It was a company wanting to stream live video from a site in London for coverage of the Royal Wedding last year. They wanted (rather adventurously) to do it on ADSL lines, needing something like 8Mb/s uplink. We managed to get the lines in with a FireBrick doing bonding so that they had around 9Mb/s upload, and all in time for the event. They streamed videos on the day. They even paid the bills (which is not cheap for 4 PSTN lines with ADSL2+ and annex M).

What made it even more of a story is that our favourite telco messed up the records on two of the lines which meant they could not accept an order for annex M (faster upload) as they did not have the line length details. It took Shaun a hell of a lot of work to get that sorted, and he knew he was working against the clock and he managed it. There are logs of him chasing our favourite telco at 2am in some cases. It was a real case of A&A staff going the extra mile, well over and above the contractual requirements to help a customer. All of the staff involved kept them informed the whole time and handled numerous questions and changes to dates and messing around very professionally and politely.

I was pleased, and the customer said they were happy with the service even.

Then we get a dispute from them - they think that they should not have to pay the install price for the lines because of the delay getting annex M on them.

What? That makes no sense. We installed the lines and got annex M and all in time. But they say they needed time to demonstrate to potential customers, and as they did not have time then did not sell as much streaming as they expected. It does not make a lot of sense as they say each stream was 2M, so they could have demonstrated as soon as they got the first 2 lines in long before the event. The delay was getting the final two annex M upgrades. Even so, them not getting customers is hardly our fault as we didn't agree an install date. In fact we make it a clear and explicit part of the contract that we don't guarantee an install date.

We explained that it made no sense. Was he saying he has losses that happened to be exactly the same as the line install costs? Anyway we were not in breach of contract, and anyway we exclude consequential losses even if we were. Sorry...

He starts spouting implied contract terms but cannot say what and where from, and that his claim is not entirely for breach of contract but cannot say what for. He even quoted Sale of Goods Act clauses relating to equipment sales, when he had not in fact complained about the equipment supplied in any way at any time, and clearly it worked as it should have. It really made no sense at all.

Even so, we did issue a good will credit for £272 which was various of the service costs before all 4 lines were working together. We did not have to, but we are nice like that.

He seemed confused by the credit. Anyway, the final email in the dispute was mine asking him to simply and clearly list exactly what he is claiming for, why, and how much. I hear no more.

Months later - The Ombudsman Service say they have a claim. The claim says we have been unhelpful and they lost business due to the delays, and states how much they have paid. It also says how much the calculate they should have paid which is £7 less. We told the Ombudsman that £7 was clearly a frivolous claim, and already more than settled, but they are going ahead anyway?!?!

Today we sent them the case file - something like 500 pages. Good luck!

ADR is Alternative Dispute Resolution. Its required for telcos like us to be part of such a scheme allowing people to take a dispute to an arbitrator instead of the county court.

This is our first case ever. No customer has taken us to ADR or court in nearly 15 years in business, and that alone is stressing me. We strive to provide good service, and to resolve disputes fairly ourselves. How can any case go as far as ADR?


The problem is, even if the arbitrator is sensible and sees we did not breach the contract and we actually bent over backwards to get the service they wanted in time in spite of serious problems (beyond our control) from our favourite telco, and as such there is no case to answer and no award... We pay £350 for that.

Yes, even a totally bonkers case and even if the arbitrator agrees it is a totally bonkers case... we pay £350.

What the hell!?!?!?!?!

We will have to see how the case pans out. In the mean time said company has paid none of their subsequent ongoing service bills, and so we are taking them to county court! Madness.


I think, certainly for business customers, we have to have a clause requiring them to pay us if they bring an invalid case to ADR. No idea if that is enforceable, but at least it is "fair". Maybe that would go to county court if I have that clause - I wonder how a county court judge would rule on the fairness of a clause forcing a "loser pays" arbitration system when, err, that is the system the country court operates... Hmmm.

P.S. had my annual diabetic review today (see other blog post) and for the first time my blood pressure was up - I was awake half the night just stressing on this whole thing being "wrong" so I am not surprised. It is not even the money - £350 is not an issue, obviously - its the injustice that stresses me.

P.P.S. I forgot the "sound bite" type paragraph for ispreview to quote...

A&A have long had concerns over the whole ADR scheme, and this case just shows how it can be abused. A clear case of A&A falling over backwards to help a customer and go way beyond the agreed contract terms, and then having a £350 bill thrown in our faces. ADR is unfair - buy definition - it is a "one-side pays regardless" arbitration scheme unlike the much cheaper country court small claims track where loser pays. We are even tempted to offer customers a scheme where we will pay their fees up front for taking us to court rather than ADR if they have a dispute, after all such a scheme would be a tenth of the cost in most cases.

2012-02-02

Two steps forward, one step back

Well, today has been interesting.

First thing was that we had a report of a /16 not routing to the Internet... The result was baffling and led to finding a rather obscure bug in BGP when using route reflectors (yes, Jon, OSPF OSPF OSPF, I know).

Basically, there are reasons to ignore a route - the RFC specifies these (cluster list showing our cluster, originid being us, etc). We do this. Good.

Sadly though we actually ignore the whole update, including the incidental withdraw prefixes in the same update. Bugger...

So upgrading around 15 boxes during the day, and I am pretty sure without losing a packet - win! - we have that fixed, and all seems fine.

Now to start seriously moving stuff over. Seems a visit to site needed - one cable showing unplugged?!?; A DSL router to install (backup management LAN); and some nice environmental sensors to install. That will be tomorrow.

DNS resolvers all working - linked in to route reflectors as local versions of our published resolvers. In fact everything now linked to two core route reflectors. Yay!

Tonight I started allowing lines to new LNSs as a test - i.e. any lines that reconnect were sent to new LNSs. We had tested a lot. We got Be, BT 20CN and BT 21CN on line and working... Good!

Then a snag - at least one wholesale L2TP customer did not route back to us on the new LNSs. Some worked, some did not. So job for tomorrow is chase them all to ensure routing all in place and allowing new LNS IP addresses through firewalls, etc. Fun!

So lines back to existing LNSs for now. If we can sort that tomorrow we can move everyone at the weekend.

We will probably set up at least one transit and one peering link on new kit tomorrow as well. Should be pretty simple and low risk (we always say that).

Still - progress...

Well, someone has to test it

I was thinking that the blog is not a bad way to explain a bit about how the network upgrade is going at A&A. We have the status pages, which are fine (well, maybe not, they need some work), but I can probably say a bit more here...

So, where are we?

The good news is that this is all just happening. The crew working on this are actually doing a good job planning and designing and, well, making it happen. They have some key deadlines they are working to, but so far everything is going pretty well. I almost feel like a director for a change, rather than an engineer. Not sure if that is scary or good.

The fun with the network this week was rather unfortunate, but yesterday we did set up the pair of route reflectors and connected almost everything up (couple of DNS servers to go). What we did do is connect the existing routers and LNS to them as well. This means we have one core network (over the old and new racks) and can start moving things.

One of the main things was testing the new link to our favourite telco. This is the primary reason for all of this - so that we can operate more than a gigabit of traffic. Initially we will be running with four gigabit fibres in to them and up to two gigabit of traffic. The load can use any of the four links in any combination which nicely allows us to run three live LNSs at well below capacity and have a fourth as a backup in case any fail. Right now we have less than a gigabit of traffic and everything can run through one LNS. The trick is to make sure that any one failure, and ideally even two failures, do no push any link or any router over capacity. The new rack can expand with more LNSs and routers to around 4 gigabit of traffic before we have to rethink things and that is probably quite a few years of expansion.

So, new host link works, and Paul has been testing his home line last night. He found and fixed a few MTU issues, but yes, it works! Well done.

But right now he as a whole rack, ten FB6000 series gigabit routers, two gigabit fibre links to the telco, and his one FTTC home line using it.

Of course, whilst this does seem like overkill, it makes no actual difference. Things go as fast as his line, as normal... We do, after all, aim not to be the bottleneck. Just amusing to think of all of that infrastructure and capacity for one line.

But it means we can do a simple LNS switch to move customers over, and get the old host link moved to the new rack so we have all four gigabit feeds. It looks like we have managed to get links via different floors (above and below us) in to the telco as well, which is good for redundancy. We also have them via different technologies (WES and EAD) so different termination kit in the rack. All of the kit in the rack is dual power and there are separate incoming power feeds.

The other good news is that this new rack also has a link for Ethernet customers. That means we can offer Ethernet via London and Maidenhead for even more redundancy. That is a link to be tested soon as well.

I'll post more on here as we make progress - but this week is key. From now on it is basically plugging things in and moving things over, and testing testing testing.

As for host names, there are a few changes. We are keeping the telco links (LNSs) as gormless (can't think of a better name), though there is a/b/c/d of them now. We are changing the edge routers from armless to aimless, an old name we used to use for edge routers, and there are a/b/c/d of them too. We have Ethernet edge routers which are core route reflectors called weightless. We still have doubtless and careless used for testing, direct L2TP and data SIMs.

The good news is that FB6000's use under 30W when running flat out, so no issues with power usage and keeping things cool. It is a very "green" network that we run.

Watch this space...

2012-02-01

Ooops

Well, what can I say - sorry to customers for the blip Tuesday evening. In fact there were a few "issues" in the late afternoon and then something of a more major "blip" lasting around 15 minutes just before 7pm.

So time to 'fess up as to what actually happened. It was us this time.

As per planned work notices we are in the middle of a major network upgrade. We have 10 shiny new routers/LNSs in the new rack and we are gradually moving things over. We want to ensure we are not the bottleneck and this means a bigger network that goes over a gigabit on various links.

One of the first steps is bringing these new routers on to our existing network. This means establishing some internal BGP links. Once this is done we can move various of the external links from one part of the network to the other in controlled steps. We are using IBGP not OSPF for various historical reasons and to date it has done what we want perfectly - we understand BGP quite well (or so we thought).

However, the main downside of IBGP is you have to mesh all of your routers. Not a problem when you have 4 of them, but when you have 10, and when connecting to the existing 4, that is a lot of BGP sessions. This is why internal routing protocols like OSPF win in such cases.

However, not a problem, we'll use route reflectors - they allow internal routing to be relayed within the internal network avoiding having to fully mesh the routers.

This is where the fun starts. Even though we have people working on this that have used BGP before joining A&A, and even though I coded the BGP in the routers myself, carefully following the RFCs, including route reflector logic, we have not actually used route reflectors in anger before.

Well, now we know - the trick is not to make a loop of route reflectors. The problem you get is a route gets injected in to this loop and then it sticks. Even if you withdraw the original announcement the loop sees its own copy (reflection) of that announcement from another route reflector and so keeps announcing it. They also tell your edge routers about the route!

To add to the fun, if you have anything not set up right in setting the next hop, you can end up with routes that go to places that don't know what to do with them (black holes).

Re-reading the RFCs this is actually quite simple, and the next test will follow the guidelines somewhat better. We will not have a loop of route reflectors but a pair of them, and the edge routers will be normal IBGP to them. This will allow the redundancy, simplicity and scalability that we want. We have fixed the next hop set up as well.

To be honest, this was a silly mistake, and one we won't make again. The impact was the minor issues in the afternoon. The actual issues were very hard to pin down as they meant some routes were broken and some were iffy (taking the wrong path in some way) but over all traffic levels stayed the same so clearly not a major issue generally. Thanks to the customers that reported what they were seeing.

Then we come to the bigger outage of 15 minutes or so. This was part of simply tidying up after the earlier problems and making the routing configs consistent. Again, a very low risk activity. We are still trying to get to the bottom of that though as it should not have caused an issue. The fix was "have you tried turning it off and then back on again" in that we reset the LNS completely, clearing all of the BGP sessions and starting from scratch. One of the jobs we still have to do this morning is trawl the logs to find why things got messed up. I would love to have spent more time tracking the problem as it happened, but getting things workings was somewhat more important. The symptoms were damn strange as sessions appeared to start up but have issues with RADIUS, even though RADIUS was apparently working and there was no apparent reasons for the sessions to have gone down in the first place. The reset meant we lost graphs for the day, which is always a nuisance.

Anyway, today's job will be carefully planning the next stages and deploying them very carefully and slowly.

Of course, and I am sure some customers will be asking this, why the hell is this not done at 3am on a Sunday morning or something? Well, yes, if this was work that was going to take out service, it would be. This is, however, work that should not actually impact service at all - it is very routine low risk stuff. It is also a case where the impact of something not being right is hard to see. If we had done this over night it would not be until 9am when a few customers say there is some "odd routing" that we would find this issue - everything looked fine when we did it!

In general we find between 5pm and 6pm to be a good time for some of this "at risk" work as it is after most business customers are finished (not all, we know), but is before the home users start (mostly) and at a time when people are still around to tell us if they can see anything not quite right.

Over night work is ideally suited to cases that take out part of the network - where the work is simple mechanical stuff - moving cables and the like - where those working on it can see they have done the job right immediately and there is nothing new. Telco work that takes out network links is scheduled for over night for this very reason.

The end result of all of this will be much more capacity in our network, and some major increases in bandwidth to our favourite telco... So sorry for the inconvenience, and we really will try not to break it like this again. Thanks for your patience.

P.S. it does seem odd not blaming our favourite telco for something. After all, over the last week we have seen BRASs reset and take out services for hundreds of customers for similar periods, but we are all kind of used to that...

P.P.S. A simple loop of route reflectors is not enough to break things - you need an ordinary IBGP link in between to lose the cluster ID.

Clocks

Some time geeks (should I say Time Lords) checked out my clocks. Seems they are impressed, saying sub microsecond. I have spent all day tryi...